Now Kubernetes will perform liveness checks on mon, mds and osd daemons.
The command will:
* call the socket (check for existence)
* execute a command and check the return code (success if 0)
This handles the case where the daemon is stuck locally and
unresponsive. It's unlikely but not impossible.
These checks bring more robustness to the implementation.
rbd-mirror and nfs have been leftover for the following reason. The
rbd-mirror socket name is different from other daemons (could be fixed
though): /run/ceph/ceph-client.rbd-mirror.a.1.94362516231272.asok also,
the command to call would need to be changed from "status" to "rbd
mirror status" so we can keep this for a later.
The nfs ganesha has no socket only a PID file which doesn't mean much.
No PID means the process does not run so Kubernetes will already handle
this and the pod will crash loop.
Signed-off-by: Sébastien Han <seb@redhat.com>
In order to ensure proper clean up of all the rook-ceph data when the cluster is deleted, we need to clean up the dataDirHostPath (var/lib/rook)
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
Current YugabyteDB operator code creates Master and TServer pods without any resource requests/limits (specifically CPU and memory).
This causes the operator to run into soft/hard memory limit issue. The fix adds recommended resource requests and limits as defaults to the pods it creates.
Closes: https://github.com/yugabyte/yugabyte-db/issues/3884
Signed-off-by: Sameer Kulkarni <samkulkarni20@gmail.com>
As of Octopus, Ceph will prevent you from creating a pool with a
replica size of 1. Allowing such pool could lead to data loss, so enable
the new option: requireSafeReplicaSize: false if you are **ABSOLUTELY**
certain that is what you want.
Closes: https://github.com/rook/rook/issues/4889
Signed-off-by: Sébastien Han <seb@redhat.com>
When the OSD on PVC is backed by a metadata block PVC, the Ceph CRUSH
device class should be set to something else rather than the rotational
property of the drive.
Closes: https://github.com/rook/rook/issues/4881
Signed-off-by: Sébastien Han <seb@redhat.com>
When running on PVC, if the backend resizes the block, Rook will expand
it and the size of the OSD will grow.
Signed-off-by: Sébastien Han <seb@redhat.com>
We now support the addition of the PVC that acts as a metadata device
for a given OSD.
For this, you need to create a new `volumeClaimTemplates`, its name must
be "metadata" otherwise, Rook won't pick it up.
A template will look like this:
```
volumeClaimTemplates:
- metadata:
name: data
spec:
resources:
requests:
storage: 10Gi
# IMPORTANT: Change the storage class depending on your environment (e.g. local-storage, gp2)
storageClassName: gp2
volumeMode: Block
accessModes:
- ReadWriteOnce
- metadata:
name: metadata
spec:
resources:
requests:
storage: 6Gi
# IMPORTANT: Change the storage class depending on your environment (e.g. local-storage, gp2)
storageClassName: gp2
volumeMode: Block
accessModes:
- ReadWriteOnce
```
We now map block and block.db directly inside the container instead of
running ceph-volume activate. This is much cleaner.
Closes: https://github.com/rook/rook/issues/3852
Signed-off-by: Sébastien Han <seb@redhat.com>
Multiple things:
1. We removed all the function/methods/tests that were used to
create and manage rook legacy OSDS as well as bringing support to
Bluestore OSD only.
It also fixes various go-lint issues in the respectives files.
2. use c-v inventory to detect available devices:
Now we rely on the 'ceph-volume inventory' command to tell us if a
device is available or not.
3. implement raw mode for osd on pvc
When an OSD will be bootstrap on a PVC, the new c-v raw mode will be
used. It consists of putting block, db and wal under the same device.
Here LVM is out of the picture and the raw device is used as is. The
implementation is backward compatible so existing OSD on PVC will LVM
will continue to operate.
Closes: https://github.com/rook/rook/issues/4363
Signed-off-by: Sébastien Han <seb@redhat.com>
The minio operator has not had community support nor
any updates since being added to Rook. Support is being
removed from Rook due to this lack of community interest.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the release of K8s 1.17 Rook needs to test on this new version.
The integration tests will now run the tests across K8s 1.13-1.17.
The pattern has been to run the tests across the most recent five
versions.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
A new CRD option `continueUpgradeAfterChecksEvenIfNotHealthy` is added.
When upgrading, Rook goes OSD by OSD and then waits for PGs to be clean
before proceeding to the next OSD. Currently, Rook waits for 5 hours but
there might be circumstances where PGs need more time to settle.
Thus setting `continueUpgradeAfterChecksEvenIfNotHealthy` to true will
pursue the upgrade process, even if PGs are not 100% active+clean.
Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1786029
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit adds status field for ceph related CRs which will be useful for knowing the ceph component status without checking the logs.
Signed-off-by: Ashish Ranjan <ashishranjan738@gmail.com>
Reduce from 5 minutes to less than 50 seconds the time needed to restart a
mgr and osds(PVCs based),rgw,mds,rbd,nfs and toolbox pod that was running
in a k8s "NotReady" Node.
A new Rook Operator setting has been added to allow the user to set the
time used in pod's toleration <node.kubernetes.io/unreachable>:
ROOK_UNREACHABLE_NODE_TOLERATION_SECONDS = <5>
New setting added in release notes
Addressed @travisn,@leseb, and @blaineEXE suggestions
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
gp2 devices present themselves as SSDs,
but have much more variable performance.
Ceph attempts to adjust some settings based on the drive type,
HDD/SSD
In this case client I/O can be impacted too much by assuming gp2 devices are like other SSDs
(an example is https://bugzilla.redhat.com/show_bug.cgi?id=1765923#c6)
In order to prevent this, rook can set the following for OSDs
if the new CR setting `tuneSlowDeviceClass` is set to `true`:
```
osd_recovery_sleep = 0.1
osd_snap_trim_sleep = 2
osd_delete_sleep = 2
```
Closes: https://github.com/rook/rook/issues/4298 and https://bugzilla.redhat.com/show_bug.cgi?id=1771047
Signed-off-by: Sébastien Han <seb@redhat.com>
Removal of OSDs is a sensitive area. If you get it wrong you could have data
loss. Now a help topic is provided for manual steps to remove OSDs when
absolutely necessary.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Removing an OSD unintentionally can result in data loss,
which is to be prevented at all costs. Rook's functionality
to remove OSDs automatically when a node was removed from the
cluster CR has led on several occasions to users having OSDs
unintentionally removed from their cluster.
If the admin decides they really need to remove OSDs they should
do it as a one-time action that they intentionally run
rather than the operator attempting to manage the deletion
with desired state. The risk of misinterpreting the desired
state, even with a typo, is too great.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Add owner reference to the discover config map so that when the
discover pod goes away we also remove its configmap.
Signed-off-by: Sébastien Han <seb@redhat.com>
As performed with mon, gateways from same ceph object store
must not run on the same host if hostNetwork is set to true
Signed-off-by: n.fraison <n.fraison@criteo.com>
Added default regex values.
Created an environment variable and using that variable the user can give custom regex values
Updated the PendingReleaseNotes.md
Signed-off-by: Nizamudeen <nia@redhat.com>
Just as with the mon, mgr, mds, rgw, and rbd-mirror daemons, do not
generate a ceph.conf in an init container for filestore OSDs.
Instead use the commandline to supply all the needed arguments for
running these OSDs.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
ceph client CRD refactoring
Add info about client crd to documentation
Add tests for client controller CRD
Signed-off-by: Mateusz Los <los.mateusz@gmail.com>
"OSD on PVC" doesn't work for PV backed by LV. Fixing this problem
by the following changes.
- Rook accepts LVM disk type.
- If a LV-backed device is passed, Rook/Ceph invokes
"ceph-volume lvm prepare" with "--data vg/lv"
instead of "--data /path/to/device".
- If a LV-backed device is passed, Rook/Ceph suppresses
activation/deactivation of VG that owns this LV.
Fixes: https://github.com/rook/rook/issues/4185
Signed-off-by: dulltz <isrgnoe@gmail.com>
Co-authored-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
This patch removes the need of using both tini and rook binary to run
the main osd container.
By removing this, we know can upgrade the operator image without
restarting the OSD since the osd pod does not use the rook image
anymore.
This patch only works when not running OSD on PVC since they are certain
actions that we cannot achieve with an init container.
A new init container has been introduced to do the activation of the
OSD, the activate files are placed in a empty dir mount and shared with
the main osd process.
Now the main osd process runs the ceph-osd daemon binary directly.
This currently has a few limitation:
* only works with bluestore OSDs
* only works when not running OSD on PVC
Closes: https://github.com/rook/rook/issues/4363
Signed-off-by: Sébastien Han <seb@redhat.com>
A new controller to bootstrap ceph-crash pod on Ceph nodes running Ceph
pods only.
This implementation is unfortunately a mix of 'hacks' to make the daemon
working correclty. The ceph-crash script faced different issues:
* it's not a daemon
* it does not have any key in cephx
* it runs ceph commands under the hood which need an admin key and a
ceph.conf
On upstream Ceph, ceph-crash needs to grow better, once that happens we
will improve our implementation.
This also enhances osd provisionConfig struct with DataPathMap
Pod volumes and volume mount needs DataPathMap to perform the right
actions on log and crash dir. Exposing DataPathMap makes that possible.
Obviously, this is exposing ceph crash reports on the host as
well as pushing them into the mgr.
Basically, if a daemon fails, it'll put its core dump into
/var/lib/ceph/crash, bindmounting this dir on the host, ensures that the
crashes don't get lost when the pod dies.
Example:
```
[leseb@tarox~/go/src/github.com/rook/rook][rgw-liveprobe !] kubectl -n rook-ceph exec -ti rook-ceph-crashcollector-minikube-574858b99c-zvg4z bash
bash: warning: setlocale: LC_CTYPE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_COLLATE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_MESSAGES: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_NUMERIC: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_TIME: cannot change locale (en_US.UTF-8): No such file or directory
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ceph crash ls
[root@rook-ceph-osd-2-7644f99695-cljzh /]# pidof ceph-osd
13258 5415 5397
[root@rook-ceph-osd-2-7644f99695-cljzh /]# kill -SIGABRT 13258
[root@rook-ceph-osd-2-7644f99695-cljzh /]# ls /var/lib/ceph/crash/
2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96 posted
... wait maximum 10 min (ceph-crash scraps every 10 minutes)
... the container will exit
[root@rook-ceph-osd-2-7644f99695-cljzh /]# exit
[leseb@tarox~/go/src/github.com/rook/rook][rgw-liveprobe !]
[leseb@tarox~/go/src/github.com/rook/rook][rgw-liveprobe !] kubectl -n rook-ceph exec -ti rook-ceph-crashcollector-minikube-574858b99c-zvg4z bash
bash: warning: setlocale: LC_CTYPE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_COLLATE: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_MESSAGES: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_NUMERIC: cannot change locale (en_US.UTF-8): No such file or directory
bash: warning: setlocale: LC_TIME: cannot change locale (en_US.UTF-8): No such file or directory
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]#
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]#
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ceph crash ls
2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96 osd.1
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ls /var/lib/ceph/crash/
posted
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# ls /var/lib/ceph/crash/posted/
2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96
[root@rook-ceph-crashcollector-minikube-574858b99c-zvg4z /]# tail /var/lib/ceph/crash/posted/2019-11-12_12\:56\:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96/log
1/ 5 mgr
1/ 5 mgrc
1/ 5 dpdk
1/ 5 eventtrace
-2/-2 (syslog threshold)
-1/-1 (stderr threshold)
max_recent 10000
max_new 1000
log_file /var/lib/ceph/crash/2019-11-12_12:56:51.404109Z_39f060e1-776d-4605-8f28-85c97e53de96/log
--- end dump of recent events ---
```
Signed-off-by: Sébastien Han <seb@redhat.com>
Co-authored-by: Rohan CJ <rohantmp@gmail.com>
The topology of the cluster should be based on the node labels rather
than a setting in the cluster CR. This allows a much richer and more
dynamic topology to be configured. The location will now be ignored
if specified in the cluster CR.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the desire to run the integration tests on the five
most recent versions of K8s, 1.16 is added and 1.11
is removed from the integration test matrix.
Signed-off-by: travisn <tnielsen@redhat.com>
As with the rest of the Ceph daemons, use the new config patters for
dir-based OSDs. Do not generate a config to use in an init container;
instead use CLI flags and the mon config database.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
Log file path on host for OSDs was dataDirHostPath/log/<namespace>
instead of dataDirHostPath/<namespace>/log as with other daemons.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
A new CRD property `PreservePoolsOnDelete` has been added to Filesystem(fs) and
Object Store(os) resources in order to increase protection against data loss.
If it is set to `true`, associated pools won't be deleted when the main
resource(fs/os) is deleted. Creating again the deleted fs/os with the same name
will reuse the preserved pools.
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
This updates the Ceph mon disaster recovery guide to use the correct
paths and commands for the current Rook Ceph release.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
A job is started to detect the Ceph version. This job now allows setting
the node affinity and tolerations with the same setting that is
specified for the mon placement. No new placement spec is required, but
we will simply use the same spec that the mons use.
Signed-off-by: travisn <tnielsen@redhat.com>
Previously, "run dir" was used to place a number of file and config on
dataDirHostPath. Now, since a lot of configs and options have moved
either on the OSD's startup CLI line or remove with bluestore we don't
need to change the default.
Also, it was confusing for user and difficult to find the socket.
The is is not breaking any config in dataDirHostPath since all the
elements (keyrings) are hardcoded in the config file.
Closes: https://github.com/rook/rook/issues/3966
Signed-off-by: Sébastien Han <seb@redhat.com>
On rare occasion, the ceph checks might not be 100% correct and the user
might decide to force an upgrade anyway.
Correctly, if we are in an upgrade, we perform `ok-to-stop` check for
each daemon, if one of them fails, we retry and eventually fail.
They are corner-cases where users would like to force the upgrade anyway.
Enhance the new CR property: "skipUpgradeChecks: true".
Closes: https://github.com/rook/rook/issues/3872
Signed-off-by: Sébastien Han <seb@redhat.com>