Commit Graph
181 Commits
Author SHA1 Message Date
Sébastien Han cba9a359a0 ceph: upgrade apply osd nautilus flag
When OSDs are running on Nautilus we always disable old osd features and
aplpy the onces for Nautilus as described in the upgrade doc.
During an upgrade or the next time an orchestration will be called the
command will be applied. The command is idempotent so we can run it each
time.
This can be backported for 1.0.3

Closes: https://github.com/rook/rook/issues/2960
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-13 22:15:15 +02:00
Ashish Ranjan 5a19ab9545 ceph: enhance server to search for rookflex
Signed-off-by: Ashish Ranjan <ashishranjan738@gmail.com>

This commit enables server to search for `rookflex` binary instead of assuming it to be present in `/usr/local/bin/`.

Fixes: https://github.com/rook/rook/issues/2486
2019-06-12 23:36:09 +05:30
Travis Nielsen 6d1ccadd4b Merge pull request #3256 from noahdesu/device-hotplug-update
discover: handle false-positives observed by users
2019-06-07 08:55:00 -07:00
Noah Watkins 3966f163ed discover: handle false-positives observed by users
the only exception to a naive device list comparison had been to ignore
drive UUID information which was unreliable when a device wasn't
formatted / partitioned. however various users have reported different
type of false positives that resulted in orchestration being run
continuously due to the wrong observation that devices were changing.

this patch fixes the cases we have observed and attempts to be slightly
more conservative in the calculation.

1. the devlinks is ignored. when a device is setup for lvm, for example,
the devlinks will be updated with different paths that point to the
device in addition to its standard paths addressable by pci address.

2. in the lvm case, the "model" field and "filesystem" field may also
change.

3. we ignore devices with devlinks that contain "usb" to avoid issues
when using usb drives.

4. be smart about detecting device availability. if a device transitions
from a non-empty (or has-partitions) state to an empty (or unpartitioned)
state then orchestration is triggered. this like observing that a device
is now available (e.g. in the allDevices case). however, when a device
transistions from empty to non-empty, then this is ignored as while it
is a change, it's generally a change associated with the new consumption
of the device.

fixes: #3059
fixes: #3185
fixes: #3131

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-06-07 08:22:08 -07:00
travisn fd341ae928 ceph: set the fsgroup only on the top level of the mount
If setting the fsgroup recursively on a shared filesystem mount is
not desirable, the fsgroup capability on the flex driver
should first be disabled in the operator env vars. Now the driver
will apply the fsgroup only at the top level instead of
recursively for the entire shared filesystem.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-06-06 17:17:29 -07:00
rohan47 a380ca7dc1 fixed ceph pg dump pgs_brief json deserialization
Signed-off-by: rohan47 <rohgupta@redhat.com>
2019-06-06 21:54:37 +05:30
Santosh Pillai ec2e6ea888 use ceph fsid to filter ceph volume OSDs
- Updated getCephVolumeOSDs method to use fsid filter while retriving devices.
- This solves "failed to fetch mon config (--no-mon-config to skip)" error when multiple ceph clusters are running on same Node

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2019-06-05 12:05:56 +05:30
Alexander TrostandMaksim Nabokikh 1d004e5b4a ceph: add metrics for flexvolume driver (#3128)
ceph: add metrics for flexvolume driver

Co-authored-by: Maksim Nabokikh <maksim.nabokikh@flant.com>
2019-05-19 13:50:15 +02:00
Santosh Pillai 7d5afa29b8 added device class pool property
- Updated code to use deviceClass property when a pool is created for both "replicated" and "erasure code".
- Updated "ceph crush rule create-..." command to use "create-replicated" instead of "create-simple"
- Updated unit tests
- Updated (ceph-pool-crd.md) documentation to reflect the changes.
- Updated pending release notes.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2019-05-15 16:45:39 +05:30
Travis Nielsen 0a9b049e17 Merge pull request #3149 from mvollman/metadata-per-osd
ceph: provision OSDs with metadataDevice config
2019-05-10 14:21:03 -04:00
Michael Vollman a839527421 ceph: provision OSDs with metadataDevice config
provision OSDs with ceph-volume when metadtaDevice is set (#3108)
always print ceph-volume report before executing ceph-volume.

Signed-off-by: Michael Vollman <michael.b.vollman@gmail.com>
2019-05-10 11:47:15 -04:00
Maksim Nabokikh cb3b8168a0 ceph: add metrics for flexvolume driver
This allows to export persistent volume metrics.

Signed-off-by: Maksim Nabokikh <maksim.nabokikh@flant.com>
2019-05-07 18:11:42 +04:00
Noah Watkins e33eb98047 ceph: run self-signed cert creation with timeout
fixes: #2784

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-05-06 10:10:05 -07:00
Noah Watkins c2dbb93096 ceph: simplify the ceph command execution interface
this adds a structure to hold various settings related to executing ceph
cli commands, and introduces a common interface for configuring and
running such commands.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-05-06 10:10:05 -07:00
Kaushal M f48f73300f ceph: Check before copying binaries in osd pods
Before performing the copy, the OSD init container now checks
destination path exists. If the destination exists, it skips the copy.

The OSD pods copy the rook and tini binaries into a shared volume using
an init container. In restarting pods, this copy could fail with 'text
file busy'(ETXTBUSY), because the destination file in the shared volume
could be under execution. The restarting pods will enter a
'Init:CrashLoopBackOff' state after this, leaving the ceph cluster
degraded. Skipping the copy will allow the init containers to complete
successfully, allowing the pod to restart.

Fixes #2674.

Signed-off-by: Kaushal M <kshlmster@gmail.com>
2019-05-03 13:08:46 +05:30
travisn 7cdfcc395a write flex settings to config file instead of env vars
The flex settings cannot be passed to the driver with environment variables.
There is no context available for returning the flex settings except
that the flex driver will look in a config file in the same directory.
These settings must be valid or else the driver will return the default
settings.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-30 14:50:43 -06:00
travisn 20f6af8a4d remove obsolete flex driver version checks
The flex driver no longer needs to check whether a kubelet
restart is needed or whether the namespace-specific driver
can be supported. The copy of k8s 1.13 types can also
be removed since we have updated to a newer k8s client.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-30 14:50:33 -06:00
Travis Nielsen 535952d3e2 Merge pull request #2777 from InfuseAI/feature/ceph-volume-metadatadevice
Support metadataDevice for ceph-volume based osd
2019-04-29 12:56:41 -06:00
Travis Nielsen 5f466b903e Merge pull request #3065 from noahdesu/hotplug-false-postives
ceph: check for equiv device config maps in operator
2019-04-26 12:03:04 -06:00
Noah Watkins fbc64ff8f4 ceph: check for equiv device config maps in operator
this is an attempt to reduce the number of false positive config map
updates used to trigger orchestration for hotplugging.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-26 09:37:25 -07:00
travisn b6568b07b4 ceph: stop disabling scrubbing during osd orchestration
Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-26 09:35:50 -06:00
Sébastien Han 5f16f8e27f ceph: fix upgrade for monitor listening on 6790 port
This commits fixes the upgrade for cluster running on 0.9.x and thus
have monitors listening on 6790.

The only downside is that the new messengers 2 protocol can not be
enabled because Ceph expects monitors to listen on 6789.
In order to enable messengers 2 your monitors will have to failover so
they can pick up both 6789 and 3300 ports, then and only then messenger
2 can be enabled.

Fixes: https://github.com/rook/rook/issues/3048 and
https://github.com/rook/rook/issues/3049
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-04-25 16:44:38 +02:00
Travis Nielsen b70bec1a76 Merge pull request #3036 from rohantmp/revertMainMon
Revert "Merge pull request #2986 from rohantmp/mainMon"
2019-04-24 15:02:58 -06:00
Travis Nielsen d11466fced Merge pull request #2900 from mackIOConsulting/fix-mgr-settings
Fix mgr settings
2019-04-24 14:44:20 -06:00
Rohan CJ fb4d23345e Revert "Merge pull request #2986 from rohantmp/mainMon"
This reverts commit b7a564e929, reversing
changes made to 6996571b5c.

Reverting: Linking ceph noout to a node taint.
The taint is applied for reasons other than maintenance as well.
This also needs a way to respect user-set noout.
Need to revisit this with a better design.

Signed-off-by: Rohan CJ <rohantmp@gmail.com>
2019-04-24 22:38:38 +05:30
Sébastien Han ab439f6602 Ceph: fix upgrade to Nautilus
This commits allows the upgrade from Mimic to Nautilus to work by:

* removing the v2 brackets on the operator ceph config file generation,
we stick with v1 and can revert this back once we deprecate mimic there
is no rush since mons keep on listening to v1.

* enable messengers 2 when the cluster runs on Nautilus

Fixes: https://github.com/rook/rook/issues/2973
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-04-24 10:57:29 +02:00
Maximilian Mack bdbd67fd7f ceph: fix mgr - dashboard settings
This fix passes on the mgrName to the configureDashboard function.

Signed-off-by: Maximilian Mack <max@mack.io>
2019-04-24 08:26:22 +02:00
Noah Watkins 16a7d6aef6 ceph: blacklist rbd devices from hotplug detection
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-23 13:44:30 -07:00
Ash WuandChia-liang Kao 118a1383e6 Support metadataDevice & databaseSizeMB for ceph-volume based osd
1. Add support of `metadataDevice` & `databaseSizeMB` options to
   OSDs provisioned by `ceph-volume`.

2. Add `ceph-volume --report` output before the actual run to
   provide more info.

Signed-off-by: Ash Wu <hSATAC@gmail.com>
Co-authored-by: Ash Wu <hSATAC@gmail.com>
Co-authored-by: Chia-liang Kao <clkao@clkao.org>
2019-04-23 16:04:21 +08:00
travisn bc87b440fb update to the k8s 1.14 client libraries
Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-18 07:51:41 -06:00
travisn 91e9a5cd04 reference K8s1.13 and client-go 1.10 and the external provisioner for flex
Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-18 07:51:41 -06:00
Rohan CJ 894f4bffbd - Set noout on cluster when storage nodes are in maintenance (unschedulable)
- Add noout functions to osd client
- export the functions DiscoverStorageNodes,GetAllStorageNodes
  from pkg/operator/ceph/cluster/osd
- Update PendingReleaseNotes.md

Signed-off-by: Rohan CJ <rohantmp@gmail.com>
2019-04-17 19:31:55 +05:30
Travis Nielsen 1562fdd7d9 Merge pull request #2946 from noahdesu/hotpluggin
ceph udev monitor / hotpluggin
2019-04-16 21:31:54 -06:00
travisn aae4d7c7f6 osds: write ceph-volume log entries as they occur
When ceph-volume is called, the default python logging will buffer everything until
the process exits, which means the pod log doesnt show what is happening
until ceph-volume is done. This change will allow logs to be written
immediately as tasks are performed inside ceph-volume

Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-16 15:47:05 -06:00
Noah Watkins e97a746386 discover: trigger device probe on udev event
this patch monitors udev events from the block subsystem via the udevadm
tool. it watches for add and delete events within a specified period
(e.g. 2 seconds) and then emits a trigger which starts a new device
probe operation.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-16 11:29:09 -07:00
Noah Watkins e0c1f5af1a discover: smarter device list change calculation
this patch ignores the UUID when determining if a device probe result
has changed. it has been observed that on some machines the UUID for the
devices are not stable. the UUID is still reported as normal through the
configmap.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-16 11:28:41 -07:00
Sébastien Han 4438672e47 ceph: expose ceph version to osd pod
We must expose the ceph version to the osd pod so that we can generate
the proper ceph configuration file. If we don't do this all the version
check will assume they run the opposite of the desired version and the
formatting of the `ceph.conf` will be wrong.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-04-15 15:57:23 +02:00
Sébastien Han b8874d7f3f ceph: allow logging control
We now run all the daemon with a new option from Nautilus 14.2.1 which
allows us to tell to a daemon to not log on file. However, we can decide
to activate logging by editing the configuration flag 'log_to_file' via
the centralized config store like this for a particular daemon:

ceph config set mon.a log_to_file true

This is useful when a daemon keeps crashing and we want to collect log
files on the system.

Fixes: https://github.com/rook/rook/issues/2881
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-04-15 10:52:03 +02:00
Alexander Trost 30589000e3 Merge pull request #2976 from galexrt/clustername_param
rookflex: Use `clusterName` if available for backwards compatibility
2019-04-13 17:56:05 +02:00
Alexander Trost 85fa159038 rookflex: Use clusterName if available for backwards compatibility
This fixes issues for users trying to delete volumes and failing because
they only have `clusterNames` set on them in Rook earlier v0.8.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2019-04-12 21:26:16 +02:00
Ramana Raja 7800d33171 ceph filesytem: use fs fail to bring down filesystem
A ceph filesystem was brought down by marking the filesystem
as down and individually failing the MDSes. From nautilus onwards,
this could be more simply and efficiently done by using
`fs fail` command. See,
https://github.com/ceph/ceph/commit/4c49f165ec1

Resolves: #2841

Signed-off-by: Ramana Raja <rraja@redhat.com>
2019-04-12 20:42:48 +05:30
travisn 5ebb651e37 ceph: update the cephcluster custom resource with ceph health
The CephCluster resource will now expose the health of the ceph cluster
so admins can query the status with k8s api or kubectl instead of
needing to run the rook toolbox. The operator will query the
ceph status periodically and update the status in the custom resource.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-10 14:41:18 -06:00
Blaine Gardner a34fbaa594 ceph: update auth caps if get/create key fails
Update integration test suite fails on an error with getting the keyring
for the mgr. The mgr's cephx capabilities have changed and need to be
updated before get, create, or get-or-create can be called. Since other
daemons' caps may change in the future, make an attempt to update the
user's caps on any key generation. This makes the key generation more
declarative from within the operator code for individual Ceph daemons.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-04-03 15:54:00 -06:00
travisn 531da9647c ceph: rely on ceph image creation success instead of rbd get info
Nautilus has an issue targeting Luminous when running the rbd info command.
Rook already has the info it needs for the call and can return it
after the rbd create command succeeds.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-03-20 15:15:36 -06:00
Noah Watkins 379cb98689 ceph: replace release name with fine-grained version
The basic idea here is to replace the release name in the ClusterInfo
object with the fine grained version information queried at runtime.
This patch also removes the Name field from the cluster spec, which was
also runtime determined and was redundant with ClusterInfo relase name.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-03-18 12:57:29 -07:00
Noah Watkins 5c9bb7d9d7 ceph: add helper to execute cmd with retries
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-03-14 14:50:54 -07:00
Ramana Raja 828d2dd366 ceph filesystem: don't use mds deactivate from mimic onwards
From ceph mimic (13.x) onwards, `mds deactivate` command is
deprecated. To scale down the number of active MDSes of a filesystem,
it's sufficient to set the desired number of active MDSes using
'max_mds' setting, and is no longer required to individually
deactivate the unwanted MDSes.

Resolves: #2725
Signed-off-by: Ramana Raja <rraja@redhat.com>
2019-03-12 15:04:11 +05:30
Travis Nielsen 4dfdf2f022 Merge pull request #2310 from SUSE/ceph-add-extra-integration-test-logging
Ceph integration tests: add more verbose logging
2019-03-07 14:40:37 -07:00
Travis Nielsen dcd24d3cca Merge pull request #2738 from SUSE/rgw-all-in-operator
Ceph rgw: configure completely from operator
2019-03-07 14:35:28 -07:00
Blaine Gardner 086fa8231c rgw: configure entirely in operator
Configure the Ceph rgw daemon completely from the operator a la the
recent changes to the Ceph mon, mgr, and mds operators.

Create the rgw deployment or daemonset first, and then create the
keyring secret for the object store with its owner reference as the
corresponding deployment or daemonset. When the replication controller
is deleted, the secret is also deleted.

The RGW's mime.types file is now stored in a configmap with a different
file created for each object store. This is primarily just a means to
get the mime.types file into the rgw pod, but the added benefit is that
the administrator can modify the configmap, which could reduce
susceptibility to file type execution vulnerabilities (worst case).

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-03-07 08:02:42 -07:00