User Tools

Site Tools


services:paddles

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
services:paddles [2026/10/02 18:46] – [Summary] djgallowayservices:paddles [2026/10/02 18:49] (current) – djgalloway
Line 19: Line 19:
  
 <code> <code>
-oc exec -n paddles $(oc get pods -n paddles -o name | grep -v build |  head -1) -- pecan expire_jobs config.py -q 14 -r 30+oc -n paddles exec deploy/paddles -- pecan expire_jobs config.py -q 14 -r 30
 </code> </code>
  
Line 36: Line 36:
 </code> </code>
  
-==== Upgrade Paddles ==== +==== Upgrading Paddles ==== 
-The paddles API / Web UI services live in Openshift in the ''paddles'' namespace.  There is a cronjob that checks quay.io/ceph/paddles for a new image every 5 minutes.  The ImageStream will automatically deploy any new images.  See https://github.com/ceph/sepia-openshift/tree/main/paddles+Merges to paddles ''main'' deploy themselves: GitHub Actions pushes ''quay.io/ceph-infra/paddles:main'' (~10 min), the ImageStream polls Quay every 5 minutes, and the trigger annotation on the Deployment rolls out the new digest.  Alembic migrations are applied by the new pod at startup - see the paddles-db README for the migration runbook. 
 + 
 +To check what's deployed or force a roll: 
 +<code> 
 +oc -n paddles get istag paddles:main 
 +oc -n paddles rollout restart deploy/paddles 
 +</code> 
 + 
 +==== Restarting ==== 
 +<code> 
 +oc -n paddles rollout restart deploy/paddles 
 +</code> 
 + 
 +==== Getting a psql shell ==== 
 +Replicas are read-only, so find the leader first: 
 +<code> 
 +oc -n paddles-db exec -it $(oc -n paddles-db get pod -l postgres-operator.crunchydata.com/role=master -o name) -c database -- psql paddles 
 +# or: patronictl list from any inst pod 
 +</code> 
 + 
 + 
 +===== Troubleshooting ===== 
 +==== Check the logs ==== 
 +<code> 
 +oc -n paddles logs deploy/paddles 
 +# or who changed something in the namespace (console Logs view, audit tenant): 
 +{log_type="audit"} |= "paddles" | json | objectRef_namespace="paddles" 
 +</code> 
 +See [[services:loki]]. 
 + 
 +==== Pods getting killed every few minutes (exit 137) ==== 
 +Almost always the ''/jobs/?description='' query.  It does a ''LIKE %%'%X%'%%'' over the 8.4M-row jobs table and postgres walks ''ix_jobs_posted'' instead of the trigram index (3-5s per call).  Enough of them saturates the gunicorn workers, liveness fails, and the pod gets killed.  Find the caller (usually a scraper hitting pulpito-ng's job history pages) and block it at soko01. 
 + 
 +==== Config changes don't take effect ==== 
 +The image's ''container_start.sh'' always uses the **sample config baked into the image**.  The ''paddles-config'' configmap mounted at ''/etc/paddles/config.py'' is ignored unless the pod ''command:'' is overridden.  Don't burn time editing the configmap and wondering why nothing changed. 
 + 
 +==== Pods Ready but API dead ==== 
 +The readiness/startup probes only test DB connectivity, not HTTP.  Check the route from a VPN host before trusting pod status. 
 + 
 +==== Dumping the DB ==== 
 +Run pg_dump against the **primary**, not a replica - replicas cancel long dumps.  Budget ~4m to dump, ~18m to restore.
services/paddles.1790966784.txt.gz · Last modified: by djgalloway