This is an old revision of the document!
Table of Contents
LONG_RUNNING_CLUSTER
Summary
A small subset of mira systems and all of the reesi systems are used in a permanent Ceph cluster.
Here's a rundown of what this cluster stores
- teuthology run logs
- quay.ceph.io containers
- chacra.ceph.com packages
- drop.ceph.com
- Files sent via ceph-post-file
Topology
Current as of 2020/03/10. OSDs are also collocated on all MON hosts.
MONs
reesi{001..005}
MGRs
reesi{004..006}
MDSs
reesi{001..003} ???
OSD hosts
mira055
mira060
mira093
reesi{001.006}
Retired hosts
mira{019,021,049,070,087,099,116,120} had all daemons removed, OSDs, evacuated and reclaimed as testnodes in February 2020. apama were retired entirely as well.
ceph.conf
This file can be saved on your workstation so you can use it as an admin node.
Current as of 2018/04/03 21:03
[global] fsid = 28f7427e-5558-4ffd-ae1a-51ec3042759a mon_host = 172.21.6.140, 172.21.6.108, 172.21.2.201, 172.21.2.202, 172.21.2.203, 172.21.2.204, 172.21.2.205 public_network = 172.21.0.0/20 # Setting below for cephmetrics.sepia.ceph.com dashboard use - dgalloway mon_health_preluminous_compat = true # ick, we have too many pgs on this cluster. mon_max_pg_per_osd = 400 [mon] debug ms = 1 debug mon = 10 [osd] debug_ms = 1 debug_osd = 10 debug_filestore = 10 setuser_match_path = $osd_data bluestore cache size = 512000000 [mds] mds cache size = 500000 mds session timeout = 120 mds session autoclose = 600 debug mds = 4 [mgr] debug mgr = 20 debug ms = 1 [mon.mira070] public addr = 172.21.6.108
Upgrading the Cluster
As of this writing, the luminous branch is the repo defined in /etc/apt/sources.list.d/ceph.list on the LRC nodes. The Ceph docs can be followed for this procedure but, basically, update and reboot each host at a time starting with MONs, MGRs, MDSs, then OSD hosts.
MONs run out of disk space
I sadly got too small of disks for the reesi when we purchased them so they occasionally run out of space in /var/log/ceph before logrotate gets a chance to run (even though it runs 4x a day. The process below will get you back up and running again but will wipe out all logs.
ansible -m shell -a "sudo /bin/sh -c 'rm -vf /var/log/ceph/*/ceph*.gz'" reesi* ansible -m shell -a "sudo /bin/sh -c 'logrotate -f /etc/logrotate.d/ceph-*'" reesi*
Replace LRC Host's root drive
On non-mon hosts
ceph osd set noouton admin hostceph osd set noscrub; ceph osd set nodeep-scrubto avoid unnecessary I/O
- Stop ceph services on OSD host
stop ceph-osd-allon Ubuntuservice ceph stop osd.#on RHEL
- Back up
/etc/cephscp root@mira###.front.sepia.ceph.com:/etc/ceph/ceph.conf .
umount /var/lib/ceph/osd/*- Back up
/var/lib/ceph/osdscp -r root@mira###.front.sepia.ceph.com:/var/lib/ceph/osd/ .
- Reimage the machine
- Install ceph packages
- If needed. set up repo file
- Also if needed, import repo GPG key
wget -qO - http://download.ceph.com/keys/release.asc | sudo apt-key add - apt-get install ceph ceph-base ceph-common ceph-osd ceph-test libcephfs1 python-cephfs ceph-deploy
- Make sure ntpd is configured and enabled
- Manually run
ntpdate $ntpserverfor one-time sync
- Configure or disable firewall
- Replace
/etc/cephand/var/lib/ceph/osdstructuresscp ceph.conf root@mira###.front.sepia.ceph.com:/etc/ceph/scp -r osd/* root@mira###.front.sepia.ceph.com:/var/lib/ceph/osd/
- Set permissions
chown -R ceph:ceph /var/lib/ceph/osd/chown ceph:ceph /etc/ceph/ceph.conf
- Create an ssh key, copy the pubkey to
/root/.ssh/authorized_keyson a monhost and runceph-deploy gatherkeys $monwhere$monis a mon host - Copy keys to their appropriate places
- For the bootstrap key,
mv ceph.bootstrap-osd.keyring /var/lib/ceph/bootstrap-osd/ceph.keyringmv ceph.client.admin.keyring /etc/ceph/chown ceph:ceph /var/lib/ceph/bootstrap-osd/ceph.keyring
reboot- Unset flags from step 1
Add blank disk as OSD
disk=sdX
ceph-disk zap /dev/$disk
ceph-disk prepare /dev/$disk
ceph-disk activate /dev/${disk}1
Replace Failing OSD disk
Evacuating OSD data
If the disk is still relatively healthy and you think it can survive a while longer, you should evacuate the data off it slowly.
- On a mon node,
ceph osd reweight $osdnum 0.75or -0.25 the current weight - Wait until recovery I/O is done and keep doing this until the OSD is reweighted to 0
Taking the OSD out of the cluster
- On a mon node,
ceph osd out $id. This makes sure there are 3 replicas of each PG evacuated.- If any recovery I/O occurs, wait for it to finish
- On the OSD host,
stop ceph-osd id=$id- Some recovery I/O will occur. This is just the cluster rebalancing. It's fine.
- Back on the mon host,
ceph osd crush remove osd.$id ceph osd down osd.$id # may not be needed as long as osd service is stopped ceph osd rm osd.$id ceph auth del osd.$id
- Unmount the disk from the OSD host
umount /var/lib/ceph/osd/ceph-$idrm -rf /var/lib/ceph/osd/ceph-$id
- Replace the disk
- On the OSD host,
disk=sdX ceph-disk zap /dev/$disk ceph-disk prepare /dev/$disk mkdir /mnt/tmp mount /dev/${disk}1 /mnt/tmp mkdir /var/lib/ceph/osd/ceph-$(cat /mnt/tmp/whoami) chown ceph:ceph /var/lib/ceph/osd/ceph-$(cat /mnt/tmp/whoami) umount /mnt/tmp ceph-disk activate /dev/${disk}1
One-liners
Most of the stuff above is no longer valuable since Ceph has evolved over time. Here's some one-liners that were useful at the time I posted them.
Restart mon service
systemctl restart ceph-28f7427e-5558-4ffd-ae1a-51ec3042759a@mon.$(hostname -s).service
Watch logs for a mon
podman logs -f $(podman ps | grep "\-mon" | awk '{ print $1 }')
