User Tools

Site Tools


services:paddles

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
services:paddles [2023/08/21 22:02] – [Summary] zmcservices:paddles [2026/10/02 18:49] (current) – djgalloway
Line 3: Line 3:
 paddles is a very simple JSON-based API used to report teuthology test results.  See https://github.com/ceph/paddles paddles is a very simple JSON-based API used to report teuthology test results.  See https://github.com/ceph/paddles
  
-The service runs on a baremetal host, [[hardware:infrastructure#pulpitofrontsepiacephcom|pulpito.front.sepia.ceph.com]], deployed via the ceph-cm-ansible role: https://github.com/ceph/ceph-cm-ansible/tree/main/roles/paddles+It runs in OpenShift in the ''paddles'' namespace.  The database is a Crunchy PGO postgres cluster in the ''paddles-db'' namespace; the app connects through the ''paddles-db-pgbouncer'' service.  See https://github.com/ceph/sepia-openshift/tree/main/paddles
  
 +paddles is **not** publicly exposed.  ''paddles.front.sepia.ceph.com'' is VPN-only.  External read-only consumers get ''%%https://pulpito.ceph.com/_paddles/runs/...%%'' (GET-only, IP-allowlisted nginx proxy on soko01).
 ===== Backups ===== ===== Backups =====
 The 'paddles' db is backed up daily by the [[services:backups#backupsh|backup.sh]] script on gitbuilder-archive. The 'paddles' db is backed up daily by the [[services:backups#backupsh|backup.sh]] script on gitbuilder-archive.
Line 11: Line 12:
  
 ===== Admin Tasks ===== ===== Admin Tasks =====
-==== Starting/Restarting service ==== 
-<code> 
-ssh ubuntu@paddles.front.sepia.ceph.com 
-sudo supervisorctl stop|stop|restart paddles 
-</code> 
  
 ==== Updating/Fixing Zombie Jobs ==== ==== Updating/Fixing Zombie Jobs ====
Line 23: Line 19:
  
 <code> <code>
-ssh paddles.front.sepia.ceph.com +oc -n paddles exec deploy/paddles -- pecan expire_jobs config.py -q 14 -r 30
-sudo docker exec -it $(sudo docker ps | grep paddles | head -n 1 | awk '{ print $1 }') sh -c "pecan expire_jobs config.py -q 14 -r 30"+
 </code> </code>
  
 ==== Adding testnodes to the inventory/DB ==== ==== Adding testnodes to the inventory/DB ====
-From your workstation,+From your workstation or anywhere with a valid ''teuthology.yaml'',
  
 <code> <code>
Line 41: Line 36:
 </code> </code>
  
-==== Upgrade Paddles ====+==== Upgrading Paddles ==== 
 +Merges to paddles ''main'' deploy themselves: GitHub Actions pushes ''quay.io/ceph-infra/paddles:main'' (~10 min), the ImageStream polls Quay every 5 minutes, and the trigger annotation on the Deployment rolls out the new digest.  Alembic migrations are applied by the new pod at startup - see the paddles-db README for the migration runbook. 
 + 
 +To check what's deployed or force a roll:
 <code> <code>
-ssh pulpito.front.sepia.ceph.com +oc -n paddles get istag paddles:main 
-sudo docker service update --image quay.io/ceph-infra/paddles:main paddles --force+oc -n paddles rollout restart deploy/paddles
 </code> </code>
 +
 +==== Restarting ====
 +<code>
 +oc -n paddles rollout restart deploy/paddles
 +</code>
 +
 +==== Getting a psql shell ====
 +Replicas are read-only, so find the leader first:
 +<code>
 +oc -n paddles-db exec -it $(oc -n paddles-db get pod -l postgres-operator.crunchydata.com/role=master -o name) -c database -- psql paddles
 +# or: patronictl list from any inst pod
 +</code>
 +
 +
 +===== Troubleshooting =====
 +==== Check the logs ====
 +<code>
 +oc -n paddles logs deploy/paddles
 +# or who changed something in the namespace (console Logs view, audit tenant):
 +{log_type="audit"} |= "paddles" | json | objectRef_namespace="paddles"
 +</code>
 +See [[services:loki]].
 +
 +==== Pods getting killed every few minutes (exit 137) ====
 +Almost always the ''/jobs/?description='' query.  It does a ''LIKE %%'%X%'%%'' over the 8.4M-row jobs table and postgres walks ''ix_jobs_posted'' instead of the trigram index (3-5s per call).  Enough of them saturates the gunicorn workers, liveness fails, and the pod gets killed.  Find the caller (usually a scraper hitting pulpito-ng's job history pages) and block it at soko01.
 +
 +==== Config changes don't take effect ====
 +The image's ''container_start.sh'' always uses the **sample config baked into the image**.  The ''paddles-config'' configmap mounted at ''/etc/paddles/config.py'' is ignored unless the pod ''command:'' is overridden.  Don't burn time editing the configmap and wondering why nothing changed.
 +
 +==== Pods Ready but API dead ====
 +The readiness/startup probes only test DB connectivity, not HTTP.  Check the route from a VPN host before trusting pod status.
 +
 +==== Dumping the DB ====
 +Run pg_dump against the **primary**, not a replica - replicas cancel long dumps.  Budget ~4m to dump, ~18m to restore.
services/paddles.1692655331.txt.gz · Last modified: by zmc