tasks:scheduled-maintenance
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| tasks:scheduled-maintenance [2017/12/21 21:08] – [Table] djgalloway | tasks:scheduled-maintenance [2018/04/03 21:30] (current) – djgalloway | ||
|---|---|---|---|
| Line 5: | Line 5: | ||
| ===== Frequency/ | ===== Frequency/ | ||
| - | The PnT Labs team has scheduled monthly maintenance on the second Friday of every month. | + | The PnT Labs team has scheduled monthly maintenance on the second Friday of every month. |
| Maintenance should not disrupt testing cycles when an upstream release is imminent. | Maintenance should not disrupt testing cycles when an upstream release is imminent. | ||
| Line 13: | Line 13: | ||
| ===== Maintenance Matrix ===== | ===== Maintenance Matrix ===== | ||
| - | ^ Service | + | ^ Service |
| - | | www.ceph.com | + | | www.ceph.com |
| - | | download.ceph.com | + | | download.ceph.com |
| - | | tracker.ceph.com | + | | tracker.ceph.com |
| - | | docs.ceph.com | + | | docs.ceph.com |
| - | | chacra.ceph.com | + | | chacra.ceph.com |
| - | | chacra dev instances | + | | chacra dev instances |
| - | | shaman | + | | shaman |
| - | | {apt-mirror, | + | | {apt-mirror, |
| - | | jenkins{2}.ceph.com | + | | jenkins{2}.ceph.com |
| - | | prado.ceph.com | + | | prado.ceph.com |
| - | | git.ceph.com | + | | git.ceph.com |
| - | | teuthology VM | No | Yes | Some | Most | + | | teuthology VM | No | Yes | Some | Most | RHEV | No | Low |
| - | | pulpito.front | + | | pulpito.front |
| - | | paddles.front | + | | paddles.front |
| - | | Cobbler | + | | Cobbler |
| - | | conserver.front | + | | conserver.front |
| - | | DHCP (store01) | + | | DHCP (store01) |
| - | | DNS | Could | Yes | N/A | Yes | RHEV/ | + | | DNS | Could | Yes | N/A | Yes |
| - | | FOG | No | Yes | No | + | | FOG | No | Yes | No |
| - | | LRC | No | Could | No | + | | LRC | No | Could | No |
| - | | gw.sepia.ceph.com | + | | gw.sepia.ceph.com |
| - | | RHEV | No | Could | Yes | No | + | | RHEV | No | Could | Yes | No | Yes | Ish |
| - | | Gluster | + | | Gluster |
| + | ===== Scheduled Maintenance Plans ===== | ||
| + | ==== CI Infrastructure Procedure ==== | ||
| + | Updating the dev chacra nodes ({1..5}.chacra.ceph.com) has little chance to affect upstream teuthology testing except while the chacra service is redeployed or a host is rebooted. | ||
| + | |||
| + | - Notify ceph-devel@ | ||
| + | - Log into each Jenkins instance, **Manage Jenkins** -> **Prepare for Shutdown** | ||
| + | - Again in Jenkins, go to **Manage Jenkins** -> **Manage Plugins** | ||
| + | - Select **All** at the bottom and click **Download now and install after restart** | ||
| + | - Wait for all jobs to finish and make sure plugins are downloaded | ||
| + | - Once all jobs are completed, ssh to each Jenkins instance | ||
| + | - Updating the jenkins package or rebooting will restart the service so: | ||
| + | - '' | ||
| + | - '' | ||
| + | - '' | ||
| + | - '' | ||
| + | - Reboot the host so you're running the latest kernel | ||
| + | - **Update Slaves** | ||
| + | - ssh to each static smithi slave (smithi{119..128} | ||
| + | - ssh to each slave-{centos, | ||
| + | - Put each irvingi node in Maintenance mode under the **Hosts** tab in the [[https:// | ||
| + | - In the RHEV Web UI, highlight each irvingi host and click **Update** | ||
| + | - Bring slave-{centos, | ||
| + | - Make sure all static slaves reconnect to Jenkins | ||
| + | - **Update chacra, mita, shaman, prado** | ||
| + | - If no service redeploy is needed for chacra, shaman, or mita, just ssh to each of those hosts and '' | ||
| + | - If a redeploy is needed, see each service' | ||
| + | - Once all the other CI hosts are up to date, update each Jenkins instance: '' | ||
| + | - This should restart Jenkins but if it doesn' | ||
| + | - '' | ||
| + | - Spot check a few jobs to make sure all plugins are working properly | ||
| + | - You can check this by commenting '' | ||
| + | - Make sure the Github hooks are working (Was a job triggered? When the job finishes, does it update the status in the PR?) | ||
| + | - Make sure postbuild scripts are running when they' | ||
| + | |||
| + | ---- | ||
| + | |||
| + | ==== Public Facing Sites Procedure ==== | ||
| + | === tracker.ceph.com and www.ceph.com === | ||
| + | For the most part, these hosts' packages can be updated and hosts rebooted whenever. | ||
| + | |||
| + | == Post-update Tasks == | ||
| + | * Log in to tracker.ceph.com and modify a bug | ||
| + | * Spot check a few pages on www.ceph.com | ||
| + | * Log in to www.ceph.com if you have a login to wordpress | ||
| + | |||
| + | ---- | ||
| + | |||
| + | === docs.ceph.com === | ||
| + | As long as there isn't a [[https:// | ||
| + | |||
| + | == Post-update Tasks == | ||
| + | * Does http:// | ||
| + | * Is the host reattached as a slave in [[https:// | ||
| + | |||
| + | ---- | ||
| + | |||
| + | === download.ceph.com === | ||
| + | Rebooting this host is disruptive to upstream testing and should be part of a planned pre-announced outage to the ceph-users and ceph-devel mailing lists. | ||
| + | |||
| + | == Post-update Tasks == | ||
| + | * Does https:// | ||
| + | |||
| + | ---- | ||
| + | |||
| + | ==== Sepia Lab Procedure ==== | ||
| + | - Send a planned outage notice to sepia at lists dot ceph.com | ||
| + | - [[services: | ||
| + | - Wait for there to be no jobs running. | ||
| + | - If you don't want to wait: | ||
| + | - Ask Yuri if '' | ||
| + | - Ask individual devs if their runs can be killed | ||
| + | - Kill idle workers< | ||
| + | ssh teuthology.front.sepia.ceph.com | ||
| + | sudo su - teuthworker | ||
| + | bin/ | ||
| + | </ | ||
| + | - Once there are no jobs running and no workers ('' | ||
| + | - Update packages on teuthology.front.sepia.ceph.com | ||
| + | - '' | ||
| + | - Update and reboot: | ||
| + | - labdashboard.front.sepia.ceph.com | ||
| + | - circle.front.sepia.ceph.com | ||
| + | - cobbler.front.sepia.ceph.com | ||
| + | - conserver.front.sepia.ceph.com | ||
| + | - drop.ceph.com | ||
| + | - git.ceph.com | ||
| + | - fog.front.sepia.ceph.com | ||
| + | - ns1.front.sepia.ceph.com | ||
| + | - ns2.front.sepia.ceph.com | ||
| + | - nsupdate.front.sepia.ceph.com | ||
| + | - vpn-pub.ovh.sepia.ceph.com | ||
| + | - satellite.front.sepia.ceph.com | ||
| + | - sentry.front.sepia.ceph.com | ||
| + | - pulpito.front.sepia.ceph.com | ||
| + | - Finally, update and reboot gw.sepia.ceph.com ((See [[services: | ||
| + | |||
| + | === Sepia Lab Post Maintenance Tasks === | ||
| + | - '' | ||
| + | - Log in to [[https:// | ||
| + | - '' | ||
| + | - '' | ||
| + | - Does git.ceph.com load? Is it up to date? | ||
| + | - Log in to [[http:// | ||
| + | - '' | ||
| + | - '' | ||
| + | - '' | ||
| + | - '' | ||
| + | - '' | ||
| + | - Run ceph-cm-ansible against that host and verify it subscribes to Satellite and can yum update | ||
| + | - Does http:// | ||
| + | - Does http:// | ||
| + | - Verify all the reverse proxies in ''/ | ||
| + | |||
| + | Finally, once all post-maintenance tasks are complete,< | ||
| + | ssh teuthology.front.sepia.ceph.com | ||
| + | sudo su - teuthworker | ||
| + | bin/ | ||
| + | ^D^D | ||
| + | </ | ||
| + | |||
| + | Check how many running workers there should be in ''/ | ||
| + | |||
| + | ===== Boilerplate Outage Notices ===== | ||
| + | ==== CI ==== | ||
| + | < | ||
| + | Hi All, | ||
| + | |||
| + | A scheduled maintenance of the CI Infrastructure is planned for YYYY-MM-DD at HH:MM UTC. | ||
| + | |||
| + | We will be updating and rebooting the following hosts: | ||
| + | jenkins.ceph.com | ||
| + | 2.jenkins.ceph.com | ||
| + | chacra.ceph.com | ||
| + | {1..5}.chacra.ceph.com | ||
| + | shaman.ceph.com | ||
| + | 1.shaman.ceph.com | ||
| + | 2.shaman.ceph.com | ||
| + | |||
| + | This means: | ||
| + | - Jenkins will be paused and stop processing new jobs so PR checks will be delayed | ||
| + | - Once there are no jobs running, all hosts will be updated and rebooted | ||
| + | - Repos on chacra nodes will be temporarily unavailable | ||
| + | |||
| + | Let me know if you have any questions/ | ||
| + | |||
| + | Thanks, | ||
| + | </ | ||
| + | |||
| + | ==== Sepia Lab ==== | ||
| + | < | ||
| + | Hi All, | ||
| + | |||
| + | A scheduled maintenance of the Sepia Lab Infrastructure is planned for YYYY-MM-DD at HH:MM UTC. | ||
| + | |||
| + | We will be updating and rebooting the following hosts: | ||
| + | teuthology.front.sepia.ceph.com | ||
| + | labdashboard.front.sepia.ceph.com | ||
| + | circle.front.sepia.ceph.com | ||
| + | cobbler.front.sepia.ceph.com | ||
| + | conserver.front.sepia.ceph.com | ||
| + | fog.front.sepia.ceph.com | ||
| + | ns1.front.sepia.ceph.com | ||
| + | ns2.front.sepia.ceph.com | ||
| + | nsupdate.front.sepia.ceph.com | ||
| + | vpn-pub.ovh.sepia.ceph.com | ||
| + | satellite.front.sepia.ceph.com | ||
| + | sentry.front.sepia.ceph.com | ||
| + | pulpito.front.sepia.ceph.com | ||
| + | drop.ceph.com | ||
| + | git.ceph.com | ||
| + | gw.sepia.ceph.com | ||
| + | |||
| + | This means: | ||
| + | - teuthology workers will be instructed to die and new jobs will not be started until the maintenance is complete | ||
| + | - DNS may be temporarily unavailable | ||
| + | - All aforementioned hosts will be temporarily unavailable for a brief time | ||
| + | - Your VPN connection will need to be restarted | ||
| + | |||
| + | I will send a follow-up "all clear" e-mail as a reply to this one once the maintenance is complete. | ||
| + | |||
| + | Let me know if you have any questions/ | ||
| + | |||
| + | Thanks, | ||
| + | </ | ||
tasks/scheduled-maintenance.1513890514.txt.gz · Last modified: by djgalloway
