Commit Graph
181 Commits
Author SHA1 Message Date
df511fb58f ci: update golangci-lint to the latest version (v1.62)
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.

Additionally, it  silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
 and string format errors found by golangci-lint, while at it.

Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
2024-12-14 14:47:30 +01:00
parth-gr 01456e8683 osd: add a new devicetype label for osds
currently device class and device type
labels were clubbed with eachother
create a seperate label, as these two labels
can be seperate

Signed-off-by: parth-gr <partharora1010@gmail.com>
2024-12-05 16:56:04 +05:30
yingshanghuangqiao 3d672a726a core: fix some comments
Signed-off-by: yingshanghuangqiao <yingshanghuangqiao@foxmail.com>
2024-07-22 22:09:52 +08:00
zer0def 9b69b5cb46 rgw: correctly handle mgr-proxied rgw cli commands in multus scenarios
Due to multi-site Ceph object resources (realms, zonegroups, zones)
reconciliation CLI commands being proxied through CephCluster's Mgr pods
in Multus-based scenarios, return codes were being additionally wrapped
by K8S' command streams and weren't adequately handled in-code, causing
reconciliation of CephObject(Realm|ZoneGroup|Zone) resources to fail, as
an exit code of 0, instead of the expected syscall.ENOENT (2) was returned.

Fixes #12833

Signed-off-by: zer0def <zv0uzqbqnncivw0afjsx79jmd19shpps3r5f2vjc6yv@definedaszero.xyz>
2023-11-20 23:42:17 +01:00
Satoru Takeuchi 3e34ebeff8 osd: handle global or node-local device class configuration correctly
Rook uses global or node-local device class configuration if device-level
configurations doesn't exist. However, after introducing device-class
level resource configuration, rook set the default value of device class.
So global or node-level device class configuration has never been used
after that.

Closes: https://github.com/rook/rook/issues/11871
Closes: https://github.com/rook/rook/issues/11826

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2023-03-27 04:28:45 +00:00
parth-gr 8e317ee074 ci: update golangci-lint version as it fails for some k8s version in 1.10
Closes: https://github.com/rook/rook/issues/11896

Signed-off-by: parth-gr <paarora@redhat.com>
2023-03-16 20:59:22 +05:30
Blaine Gardner e88b920411 Merge pull request #11221 from zhucan/bugfix-11204
file: check if fs exists before checking dependencies
2023-03-01 12:18:03 -07:00
parth-grandBlaine Gardner 69ae9569cd build: update k8s version to 1.26.1
Update the Kubernetes API version used to 1.26.1, and start testing
against Kubernetes version 1.26.1 in CI.

Co-authored-by: parth-gr <paarora@redhat.com>
Co-authored-by: Blaine Gardner <blaine.gardner@redhat.com>

Signed-off-by: parth-gr <paarora@redhat.com>
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2023-02-23 20:18:47 -07:00
zhucan 46c2bef9bc file: check if fs exists before checking dependencies
Signed-off-by: zhucan <zhucan.k8s@gmail.com>
2023-02-20 11:53:59 +08:00
parth-gr a84daf9bf0 core: change io/ioutil package to use io and os package
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues

Signed-off-by: parth-gr <paarora@redhat.com>
2023-02-17 20:38:29 +05:30
Rakshith R bde286ef49 osd: re-open encrypted disk during osd-prepare-job if closed
This commit implements this corner case during osd-prepare job.
```
The encrypted block is not opened, this is an extreme corner case
The OSD deployment has been removed manually AND the node rebooted
So we need to re-open the block to re-hydrate the OSDInfo.

Handling this case would mean, writing the encryption key on a
temporary file, then call luksOpen to open the encrypted block and
then call ceph-volume to list against the opened encrypted block.
We don't implement this, yet and return an error.
```
When underlying PVC for osd are CSI provisioned, the encrypted device
is closed when PVC is unmounted due to osd pod being deleted.
Therefore, this may occur more frequently and needs to be handled.
This commit implements the fix for the same.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-11-30 18:54:23 +05:30
Shinya Hayashi 05875a3f4f osd: support loop devices for test clusters
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.

This feature is intended to be used for testing purposes only.

Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
2022-11-09 06:45:41 +00:00
Blaine Gardner 005000212c nfs: add kerberos client security support
Add support for enabling kerberos for client authentication.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-09-13 08:46:52 -06:00
Travis Nielsen 5ef8d15659 build: format comments for go 1.19
The tool gofmt in go 1.19 requires certain formatting
in the comments section for better rendering.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-11 12:41:10 -06:00
Satoru Takeuchi 7e571f6114 osd: support OSD on logical volume in host-based cluster
Rook supports raw mode OSD in host-based cluster. So we can also
support OSD on logical volume in this kind of cluster.

Logical volumes aren't picked by filters (i.e. `useAllDevices: true`
and `device{Path,}Filter` to avoid unwanted LV consumption on upgrade.

Closes: https://github.com/rook/rook/issues/2047

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-07-11 21:08:50 +00:00
Satoru Takeuchi e66b35aa55 Merge pull request #10230 from leseb/fix-9470
core: append filesystem property on disk
2022-05-10 09:07:09 +09:00
Sébastien Han 8328ec6d31 core: add mountpoint detection to the prepare pod
Amount various filters let's add a mountpoint property and skip the
device if its filesystem is mounted.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-05-09 15:26:50 +02:00
Sébastien Han 77b2e10c14 core: add more debug info to the prepare pod
When running the device property commands, let's also log their results
for better debugging/visibility.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-05-09 15:26:50 +02:00
Satoru Takeuchi 4c46dd1abd osd: fix disk uuid management
- disk UUID should only be got for "disk" type device.
- It's not necessary to fail prepare job when failing to get disk UUID
- We should consider that `sgdisk` reports UUID even if there is no GPT.

Closes: https://github.com/rook/rook/issues/9948

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-05-09 18:25:09 +09:00
subhamkrai 24802c559e core: fix golangci linter
fix golangci linter

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-04 20:59:31 +05:30
Divyansh Kamboj 9008409f87 core: add context parameter to functions
This commit adds context parameter to various functions, and remove the
usage of context.TODO.

Closes: https://github.com/rook/rook/issues/8701
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2022-03-22 08:19:07 +05:30
subhamkrai 6bdd5b2062 core: remove obsolete ceph-volume executor code
In volume.go, remove if condition in `callCephVolume`
method where we always pass false.

In exec.go, removing `ExecuteCommandWithFile*` methods
which was not used anywhere.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-01-28 20:17:55 +05:30
Sébastien Han 05775b068d osd: handle removal of encrypted osd deployment
This is handling a tricky scenario where the OSD deployment is manually
removed and the OSD never reconvers. This is unlikely to happen, but
still OSD should be able to run after that action. Essentially after a
manual deletion, we need to run the prepare job again to re-hydrate the
OSD information so that the OSD deployment can be deployed.
On encryption, it is a little bit tricky since ceph-volume list again
the main block won't return anything, so we need to target the encrypted
block to list.
There is another case this PR does not handle, which is the removal of
the OSD deployment and then the node is restarted. This means that the
encrypted container is not opened anymore. However, opening it requires
more work like writing the key on the filesystem (if not coming from the
Kubernete secret, eg,. KMS vault) and then run luksOpen. This is an
extreme corner case probably not worth worrying about for now.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-21 10:51:09 +01:00
Satoru Takeuchi 77e7c249cc core: remove unnecessary option
rook command doesn't interpret `logtostderr` option. It's OK to just
remove this option because `capnslog` outputs all logs to stdout
by default. It's better to keep `AddGoFlagSet()` call because
some libraries might define their own flags with Go's `flag` package.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-11-15 11:53:45 +00:00
Blaine Gardner 01e2feaef5 ceph: get rid of dynamic clientset
Stop using the dynamic clientset in favor of the controller-runtime
clientset.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-19 17:05:38 -06:00
Sébastien Han 1e9d24ac77 ceph: print the c-v output when inventory command fails
Hopefully, this will give us more hints when a failure occurs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-13 18:08:06 +02:00
Sébastien Han 4471f3a92e osd: do not hide errors
The previous exit 32 check for loop device is 5 years old. Also, if the
device cannot be read it will be skipped anyway so let's report the
error and not hide it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-07 17:39:58 +02:00
Blaine Gardner 7586cea049 core: create TRACE_INSECURE log level
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.

If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.

Resolves #8778

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-01 12:09:19 -06:00
Blaine Gardner 7b9293624a rgw: update period if period does not exist
Rook should update the RGW object store's period if the period doesn't
yet exist. This protects us from the case where the
'radosgw-admin period update --commit` command fails and the
CephObjectStore controller reconciles again.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-28 14:17:04 -06:00
Blaine Gardner bc494f5174 core: add missing error type check to exec
In ExtractExitCode, there is one error type that can be valid as a
pointer or not-as-a-pointer. Add a case to the type check for the
non-pointer condition.

Resolves #8280

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-22 14:16:52 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han b7f55362b2 ceph: print the output on errors
Sometimes the error does not tell much, so as `exit status 1` and
printing the output along returning the error is useful.

For instance, I saw a job failing with no osd and the prepare job had
those lines:

```
exec: Running command: lsblk /dev/sdb1 --bytes --nodeps --pairs ....
inventory: skipping device "sdb1". exit status 1
```

We need to understand more about the lsblk issue.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-09 12:42:01 +02:00
Sébastien Han d4dd0577f9 Merge pull request #8529 from sp98/id-mapping
ceph: add ClusterID and PoolID mappings between local and peer cluster
2021-08-31 17:32:14 +02:00
Santosh Pillai 3f8abec403 ceph: add ClusterID and PoolID mappings between local and peer cluster
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters

This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-08-31 20:08:56 +05:30
parth-gr b1cfeb9dbe core: removing util.set package files
We are no longer using package util.set
for creating instances,
and instead of that using sets.String package

Closes: https://github.com/rook/rook/issues/8479
Signed-off-by: parth-gr <paarora@redhat.com>
2021-08-25 13:15:14 +05:30
Sébastien Han 73b0ce08ff ceph: do not stack trace on error
We should not use %+v when printing errors and wrapping it in the
caller. Passing %+v to errors.Wrap() returns a strack trace.

See:

```
2021-08-03 09:42:40.460946 E | ceph-file-controller: failed to reconcile failed to create filesystem "myfs2": failed to start deployment for MDS "a" for filesystem "myfs2": failed to update mds deployment "rook-ceph-mds-myfs2-a": failed to check if deployment "rook-ceph-mds-myfs2-a" can continue: failed to check if we can continue the deployment rook-ceph-mds-myfs2-a: failed to check if rook-ceph-mds-myfs2-a was ok to continue: max retries exceeded, last err: mds myfs2-a is up:creating, bad state
github.com/rook/rook/pkg/daemon/ceph/client.MdsActiveOrStandbyReplay
        /home/leseb/go/src/github.com/rook/rook/pkg/daemon/ceph/client/status.go:286
github.com/rook/rook/pkg/daemon/ceph/client.okToContinueMDSDaemon.func1
        /home/leseb/go/src/github.com/rook/rook/pkg/daemon/ceph/client/upgrade.go:215
github.com/rook/rook/pkg/util.Retry
        /home/leseb/go/src/github.com/rook/rook/pkg/util/retry.go:31
github.com/rook/rook/pkg/daemon/ceph/client.okToContinueMDSDaemon
        /home/leseb/go/src/github.com/rook/rook/pkg/daemon/ceph/client/upgrade.go:214
github.com/rook/rook/pkg/daemon/ceph/client.OkToContinue
        /home/leseb/go/src/github.com/rook/rook/pkg/daemon/ceph/client/upgrade.go:181
github.com/rook/rook/pkg/operator/ceph/cluster/mon.UpdateCephDeploymentAndWait.func1
        /home/leseb/go/src/github.com/rook/rook/pkg/operator/ceph/cluster/mon/spec.go:370
github.com/rook/rook/pkg/operator/k8sutil.UpdateDeploymentAndWait
        /home/leseb/go/src/github.com/rook/rook/pkg/operator/k8sutil/deployment.go:113
github.com/rook/rook/pkg/operator/ceph/cluster/mon.UpdateCephDeploymentAndWait
        /home/leseb/go/src/github.com/rook/rook/pkg/operator/ceph/cluster/mon/spec.go:383
github.com/rook/rook/pkg/operator/ceph/file/mds.(*Cluster).startDeployment
        /home/leseb/go/src/github.com/rook/rook/pkg/operator/ceph/file/mds/mds.go:216
github.com/rook/rook/pkg/operator/ceph/file/mds.(*Cluster).Start
        /home/leseb/go/src/github.com/rook/rook/pkg/operator/ceph/file/mds/mds.go:138
github.com/rook/rook/pkg/operator/ceph/file.createFilesystem
        /home/leseb/go/src/github.com/rook/rook/pkg/operator/ceph/file/filesystem.go:81
github.com/rook/rook/pkg/operator/ceph/file.(*ReconcileCephFilesystem).reconcileCreateFilesystem
        /home/leseb/go/src/github.com/rook/rook/pkg/operator/ceph/file/controller.go:373
github.com/rook/rook/pkg/operator/ceph/file.(*ReconcileCephFilesystem).reconcile
        /home/leseb/go/src/github.com/rook/rook/pkg/operator/ceph/file/controller.go:288
github.com/rook/rook/pkg/operator/ceph/file.(*ReconcileCephFilesystem).Reconcile
        /home/leseb/go/src/github.com/rook/rook/pkg/operator/ceph/file/controller.go:167
sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).reconcileHandler
        /home/leseb/go/pkg/mod/sigs.k8s.io/controller-runtime@v0.9.0/pkg/internal/controller/controller.go:298
sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).processNextWorkItem
        /home/leseb/go/pkg/mod/sigs.k8s.io/controller-runtime@v0.9.0/pkg/internal/controller/controller.go:253
sigs.k8s.io/controller-runtime/pkg/internal/controller.(*Controller).Start.func2.2
        /home/leseb/go/pkg/mod/sigs.k8s.io/controller-runtime@v0.9.0/pkg/internal/controller/controller.go:214
runtime.goexit
        /usr/local/go/src/runtime/asm_amd64.s:1371
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-03 11:48:54 +02:00
Satoru Takeuchi 30e4fbb01f ceph: make the timeout of ceph commands cofigurable
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-27 12:42:16 +00:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Blaine Gardner c8a2db368a ceph: disable raw mode for disks
Even if we can use raw mode, do NOT use raw mode on disks. Ceph bluestore disks can
sometimes appear as though they have "phantom" Atari (AHDI) partitions created on them
when they don't in reality. This is due to a series of bugs in the Linux kernel when it
is built with Atari support enabled. This behavior does not appear for raw mode OSDs on
partitions, and we need the raw mode to create partition-based OSDs. We cannot merely
skip creating OSDs on "phantom" partitions due to a bug in `ceph-volume raw inventory`
which reports only the phantom partitions (and malformed OSD info) when they exist and
ignores the original (correct) OSDs created on the raw disk.

Resolves https://github.com/rook/rook/issues/7940

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-07-22 09:59:56 -06:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Sébastien Han 8c589a3785 ceph: silence harmless errors
Let's not look at misleading errors if the operator is still
initializing.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 15:37:08 +02:00
Blaine Gardner 037945aabc Merge pull request #8112 from BlaineEXE/dependencies-2-son-of-dependencies
ceph: block delete object store when buckets and/or users exist
2021-06-30 12:45:58 -06:00
Blaine Gardner 7f516b9e3d ceph: ignore atari partitions when scanning disks
Ceph bluestore raw disks can sometimes appear as though they have Atari
(AHDI) partitions on them. If a disk has Atari partitions, we just
ignore them as though they don't exist. This should be a safe assumption
since the hardware was last manufactured in 1992 and likely can't run
Kubernetes.

If we don't ignore the Atari partitions, Rook can create a new OSD on a
disk that is already running an OSD, corrupting the first and possibly
also the latest OSD. This can happen an arbitrary number of times per
disk.

More info: https://github.com/rook/rook/issues/7940

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-30 09:42:33 -06:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Sébastien Han da56354aa3 ceph: run CLI output with Info logging
In e8f9cfcb71, the logging was replaced by
Debug which is probably a mistake given the intent of the commit to
enable debug logging on the prepare job.
However, this code is only triggered when running a ceph-osd with the
rook binary, so we must run Info and run Debug since Debug is not
activated.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-22 16:43:47 +02:00
Blaine Gardner 21e290e003 ceph: implement dependencies for CephCluster
Implement the first step of `design/ceph/resource-dependencies.md` to
add dependency checking when deleting a CephCluster.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-10 10:06:10 -06:00
Denis Egorenko 8a556fd847 ceph: add ability to set resource limit for OSDs based on device classes
Currently it is not possible to set resource limits based on device classes
for different OSDs. Now adding an ability to use predefined keys for main
cluster spec Resource section to reflect resource limits for different OSDs
with different device classes.

Closes: https://github.com/rook/rook/issues/8007
Signed-off-by: Denis Egorenko <degorenko@mirantis.com>
2021-06-08 17:46:41 +04:00
Blaine Gardner d78b2a41d6 ceph: work around radosgw-admin fifo file io error
Work around issue https://github.com/rook/rook/issues/7573
and make sure integration tests check for regressions.

Eventually we should use the RADOS Gateway admin REST API, but for now
we need to work around an issue where the built version of
'radosgw-admin' has incompatibilities with the RADOS Gateway version
running in the Ceph cluster.

Of note, Rook built on the Ceph Pacific image will not support
some 'radosgw-admin' commands to Ceph Nautilus (v14) or Octopus (v15)
clusters.

Further complicating matters, the flag used for the workaround changes
between Ceph v16.2.0 and v16.2.1 (both Pacific).

This bug needs to be treated a little differently than most of the ways
Rook handles different commands for different Ceph versions because this
is based on the Ceph version that is installed in the container with the
Rook operator primarily and not the version of Ceph running in the
cluster.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-04-13 12:08:02 -06:00
Blaine Gardner 795124b7a8 ceph: update osds in parallel
Update OSDs in parallel per the design in
design/ceph/update-osds-in-parallel.md

The max number of OSDs updated in parallel is currently fixed at 20.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-03-29 10:55:28 -06:00
Sébastien Han 2d75135579 ceph: embed the buffer in the error
Somehow the combined output does not contain the stderr in the errorn
only in the buffer.
For example:

err = "error exit 1"
buf = "Error EBUSY: not enough monitors would be available () after stopping mons [a]"

So we now embed the buf in the error so that the caller does not need to
print the buffer as well as the error.
Also, the caller can more easily decide to ignore the error by
instropecting the string.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-02-03 17:28:47 +01:00