When OSDs are running on Nautilus we always disable old osd features and
aplpy the onces for Nautilus as described in the upgrade doc.
During an upgrade or the next time an orchestration will be called the
command will be applied. The command is idempotent so we can run it each
time.
This can be backported for 1.0.3
Closes: https://github.com/rook/rook/issues/2960
Signed-off-by: Sébastien Han <seb@redhat.com>
the only exception to a naive device list comparison had been to ignore
drive UUID information which was unreliable when a device wasn't
formatted / partitioned. however various users have reported different
type of false positives that resulted in orchestration being run
continuously due to the wrong observation that devices were changing.
this patch fixes the cases we have observed and attempts to be slightly
more conservative in the calculation.
1. the devlinks is ignored. when a device is setup for lvm, for example,
the devlinks will be updated with different paths that point to the
device in addition to its standard paths addressable by pci address.
2. in the lvm case, the "model" field and "filesystem" field may also
change.
3. we ignore devices with devlinks that contain "usb" to avoid issues
when using usb drives.
4. be smart about detecting device availability. if a device transitions
from a non-empty (or has-partitions) state to an empty (or unpartitioned)
state then orchestration is triggered. this like observing that a device
is now available (e.g. in the allDevices case). however, when a device
transistions from empty to non-empty, then this is ignored as while it
is a change, it's generally a change associated with the new consumption
of the device.
fixes: #3059fixes: #3185fixes: #3131
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
If setting the fsgroup recursively on a shared filesystem mount is
not desirable, the fsgroup capability on the flex driver
should first be disabled in the operator env vars. Now the driver
will apply the fsgroup only at the top level instead of
recursively for the entire shared filesystem.
Signed-off-by: travisn <tnielsen@redhat.com>
- Updated getCephVolumeOSDs method to use fsid filter while retriving devices.
- This solves "failed to fetch mon config (--no-mon-config to skip)" error when multiple ceph clusters are running on same Node
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
- Updated code to use deviceClass property when a pool is created for both "replicated" and "erasure code".
- Updated "ceph crush rule create-..." command to use "create-replicated" instead of "create-simple"
- Updated unit tests
- Updated (ceph-pool-crd.md) documentation to reflect the changes.
- Updated pending release notes.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
provision OSDs with ceph-volume when metadtaDevice is set (#3108)
always print ceph-volume report before executing ceph-volume.
Signed-off-by: Michael Vollman <michael.b.vollman@gmail.com>
this adds a structure to hold various settings related to executing ceph
cli commands, and introduces a common interface for configuring and
running such commands.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
Before performing the copy, the OSD init container now checks
destination path exists. If the destination exists, it skips the copy.
The OSD pods copy the rook and tini binaries into a shared volume using
an init container. In restarting pods, this copy could fail with 'text
file busy'(ETXTBUSY), because the destination file in the shared volume
could be under execution. The restarting pods will enter a
'Init:CrashLoopBackOff' state after this, leaving the ceph cluster
degraded. Skipping the copy will allow the init containers to complete
successfully, allowing the pod to restart.
Fixes#2674.
Signed-off-by: Kaushal M <kshlmster@gmail.com>
The flex settings cannot be passed to the driver with environment variables.
There is no context available for returning the flex settings except
that the flex driver will look in a config file in the same directory.
These settings must be valid or else the driver will return the default
settings.
Signed-off-by: travisn <tnielsen@redhat.com>
The flex driver no longer needs to check whether a kubelet
restart is needed or whether the namespace-specific driver
can be supported. The copy of k8s 1.13 types can also
be removed since we have updated to a newer k8s client.
Signed-off-by: travisn <tnielsen@redhat.com>
this is an attempt to reduce the number of false positive config map
updates used to trigger orchestration for hotplugging.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
This commits fixes the upgrade for cluster running on 0.9.x and thus
have monitors listening on 6790.
The only downside is that the new messengers 2 protocol can not be
enabled because Ceph expects monitors to listen on 6789.
In order to enable messengers 2 your monitors will have to failover so
they can pick up both 6789 and 3300 ports, then and only then messenger
2 can be enabled.
Fixes: https://github.com/rook/rook/issues/3048 and
https://github.com/rook/rook/issues/3049
Signed-off-by: Sébastien Han <seb@redhat.com>
This reverts commit b7a564e929, reversing
changes made to 6996571b5c.
Reverting: Linking ceph noout to a node taint.
The taint is applied for reasons other than maintenance as well.
This also needs a way to respect user-set noout.
Need to revisit this with a better design.
Signed-off-by: Rohan CJ <rohantmp@gmail.com>
This commits allows the upgrade from Mimic to Nautilus to work by:
* removing the v2 brackets on the operator ceph config file generation,
we stick with v1 and can revert this back once we deprecate mimic there
is no rush since mons keep on listening to v1.
* enable messengers 2 when the cluster runs on Nautilus
Fixes: https://github.com/rook/rook/issues/2973
Signed-off-by: Sébastien Han <seb@redhat.com>
1. Add support of `metadataDevice` & `databaseSizeMB` options to
OSDs provisioned by `ceph-volume`.
2. Add `ceph-volume --report` output before the actual run to
provide more info.
Signed-off-by: Ash Wu <hSATAC@gmail.com>
Co-authored-by: Ash Wu <hSATAC@gmail.com>
Co-authored-by: Chia-liang Kao <clkao@clkao.org>
When ceph-volume is called, the default python logging will buffer everything until
the process exits, which means the pod log doesnt show what is happening
until ceph-volume is done. This change will allow logs to be written
immediately as tasks are performed inside ceph-volume
Signed-off-by: travisn <tnielsen@redhat.com>
this patch monitors udev events from the block subsystem via the udevadm
tool. it watches for add and delete events within a specified period
(e.g. 2 seconds) and then emits a trigger which starts a new device
probe operation.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
this patch ignores the UUID when determining if a device probe result
has changed. it has been observed that on some machines the UUID for the
devices are not stable. the UUID is still reported as normal through the
configmap.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
We must expose the ceph version to the osd pod so that we can generate
the proper ceph configuration file. If we don't do this all the version
check will assume they run the opposite of the desired version and the
formatting of the `ceph.conf` will be wrong.
Signed-off-by: Sébastien Han <seb@redhat.com>
We now run all the daemon with a new option from Nautilus 14.2.1 which
allows us to tell to a daemon to not log on file. However, we can decide
to activate logging by editing the configuration flag 'log_to_file' via
the centralized config store like this for a particular daemon:
ceph config set mon.a log_to_file true
This is useful when a daemon keeps crashing and we want to collect log
files on the system.
Fixes: https://github.com/rook/rook/issues/2881
Signed-off-by: Sébastien Han <seb@redhat.com>
This fixes issues for users trying to delete volumes and failing because
they only have `clusterNames` set on them in Rook earlier v0.8.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
A ceph filesystem was brought down by marking the filesystem
as down and individually failing the MDSes. From nautilus onwards,
this could be more simply and efficiently done by using
`fs fail` command. See,
https://github.com/ceph/ceph/commit/4c49f165ec1Resolves: #2841
Signed-off-by: Ramana Raja <rraja@redhat.com>
The CephCluster resource will now expose the health of the ceph cluster
so admins can query the status with k8s api or kubectl instead of
needing to run the rook toolbox. The operator will query the
ceph status periodically and update the status in the custom resource.
Signed-off-by: travisn <tnielsen@redhat.com>
Update integration test suite fails on an error with getting the keyring
for the mgr. The mgr's cephx capabilities have changed and need to be
updated before get, create, or get-or-create can be called. Since other
daemons' caps may change in the future, make an attempt to update the
user's caps on any key generation. This makes the key generation more
declarative from within the operator code for individual Ceph daemons.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
Nautilus has an issue targeting Luminous when running the rbd info command.
Rook already has the info it needs for the call and can return it
after the rbd create command succeeds.
Signed-off-by: travisn <tnielsen@redhat.com>
The basic idea here is to replace the release name in the ClusterInfo
object with the fine grained version information queried at runtime.
This patch also removes the Name field from the cluster spec, which was
also runtime determined and was redundant with ClusterInfo relase name.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
From ceph mimic (13.x) onwards, `mds deactivate` command is
deprecated. To scale down the number of active MDSes of a filesystem,
it's sufficient to set the desired number of active MDSes using
'max_mds' setting, and is no longer required to individually
deactivate the unwanted MDSes.
Resolves: #2725
Signed-off-by: Ramana Raja <rraja@redhat.com>
Configure the Ceph rgw daemon completely from the operator a la the
recent changes to the Ceph mon, mgr, and mds operators.
Create the rgw deployment or daemonset first, and then create the
keyring secret for the object store with its owner reference as the
corresponding deployment or daemonset. When the replication controller
is deleted, the secret is also deleted.
The RGW's mime.types file is now stored in a configmap with a different
file created for each object store. This is primarily just a means to
get the mime.types file into the rgw pod, but the added benefit is that
the administrator can modify the configmap, which could reduce
susceptibility to file type execution vulnerabilities (worst case).
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>