Commit Graph
268 Commits
Author SHA1 Message Date
Travis Nielsen d55669293a mon: set stretch tiebreaker reliably during failover
The failover of the arbiter mon in a stretch cluster was sometimes
failing due to the new tiebreaker not being set in ceph.
Rook would repeatedly try to remove the old tiebreaker mon
and keep failing because the new tiebreaker had not been set.
Now we make setting the tiebreaker idempotent in case the operator
restarts in the middle of the operation or some other corner
case causes the expected tiebreaker to be set. In that case,
the next reconcile will also ensure the tiebreaker mon is
set as expected.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-02 12:50:07 -07:00
Sébastien Han d8a8b05c1f cephfs-mirror: use combined output to possibly catch peer import error
By adding a combined output to the executor we might be able to fetch more
error messages.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-26 15:53:30 +01:00
Sébastien Han 7f7a72d941 core: add the ability to execute ceph commands with a combined output
Sometimes Ceph uses a different standard output to return errors or
merges standard error to standard out. So let's allow some commands to
return both in the output.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-26 15:52:49 +01:00
Sébastien Han d1cdba420a cephfs-mirror: try to mitigate peer import error
Some users have reported issues while adding the token, this is not
always reproducable so perhaps it's a typo when importing the token and
adding trailing spaces.

Closes: https://github.com/rook/rook/issues/9151
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-26 15:45:01 +01:00
Sébastien Han 7402c2cce6 osd: check if osd is ok-to-stop before removal
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.

So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-25 10:21:32 +01:00
Sébastien Han c3accdcc3a nfs: only set the pool size when it exists
For CRD not using the new nfs spec that includes the pool settings,
applying the "size" property won't work since it is set to 0. The pool
still gets created but returns an error. The loop is re-queued but on
the second run the pool is detected so no further configuration is done.

Closes: https://github.com/rook/rook/issues/9205
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-23 11:07:43 +01:00
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Sébastien Han 10a114a45a core: fail if config dir creation fails
If the operator fails to create the operator's configuration directory
then we should fail the Operator.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-08 08:48:35 +01:00
Sébastien Han 8cb6cdcea4 rbd-mirror: use a sorted list for peer token content
The monitor list was not sorted, so each time we were reconciling, the
peer secret token will see its content updated with randomized
monitors. This would enter our predicate and trigger a reconcile.
Potentially an endless one, if the randomized list is already different.

Closes: https://github.com/rook/rook/issues/9076
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-03 17:48:54 +01:00
Sébastien Han 3b5bbf28d5 Merge pull request #9003 from leseb/vault-k8s-auth
osd: add support for k8s with vault kms
2021-10-25 18:02:48 +02:00
Sébastien Han 18a4047679 osd: add support for k8s with vault kms
Rook cluster-wide encryption can now use the native Kubernetes
authentication to interact with vault KMS instead of using the token
method.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-21 13:59:28 +02:00
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
Sébastien Han df2b6faa96 Merge pull request #8995 from leseb/no-error-merge
ceph: only merge stderr on error
2021-10-18 18:28:04 +02:00
Sébastien Han 0e41e36ade ceph: only merge stderr on error
Previously we were merging the stderr even if it was empty, leading to
unmarshall errors.
The error simulation was done here
https://play.golang.org/p/Sk2yw9GUWNu.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-18 17:18:57 +02:00
Travis Nielsen 12689bd119 ceph: enable mon failover for the arbiter in stretch mode
Prior to ceph v16.2.7 the failover of the arbiter mon was
not supported. Now the new tiebreaker mon can be set during
the failover event and provide more dynamic stability to
the mon quorum if another node is available in the arbiter
zone.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-18 08:58:25 -06:00
Sébastien Han 28cc6f5514 ceph: remove default value for pool compression
Using a default value for CompressionMode to none effectively overrides
any values for Parameters. It is deprecated but still takes precedence.
Which means that in its previous form, Parameters was always ignored
since CompressionMode was always set to none when empty.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-13 17:06:22 +02:00
Blaine Gardner 7586cea049 core: create TRACE_INSECURE log level
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.

If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.

Resolves #8778

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-01 12:09:19 -06:00
Sébastien Han 17999bcb3d ceph: do not build all the args to remote exec cmd
When proxying commands to the cmd-proxy container we don't need to build
the command line with the same flags as the operator. The cmd-proxy
container does not use any ceph config file and just relies on the
CEPH_ARGS environment variable in the container. So passing the same
args as the operator causes to fail since we don't have a ceph config
file in `/var/lib/rook/openshift-storage/openshift-storage.config` thus
the remote exec fails with:

```
global_init: unable to open config file from search list ...
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-28 19:00:32 +02:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han d4dd0577f9 Merge pull request #8529 from sp98/id-mapping
ceph: add ClusterID and PoolID mappings between local and peer cluster
2021-08-31 17:32:14 +02:00
Santosh Pillai 3f8abec403 ceph: add ClusterID and PoolID mappings between local and peer cluster
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters

This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-08-31 20:08:56 +05:30
Sébastien Han 9744f0f2c6 Merge pull request #8583 from leseb/cancel-retry
ceph: do not check ok-to-stop when OSDs are in CLBO
2021-08-25 17:38:06 +02:00
Sébastien Han 88048cc4e8 ceph: do not check ok-to-stop when OSDs are in CLBO
This handles the scenario where the OSDs have been created but not yet
started due to a wrong CR configuration.
For instance, when OSDs are encrypted and Vault is used to store
encryption keys, if the KV version is incorrect during the cluster
initialization the OSDs will fail to start and stay in CLBO until the
CR is updated again with the correct KV version so that it can start.

For this scenario, if the CRUSH map has no host registered yet it's fair
to assume the initialization broke and we need to fix it. So when don't
need to call ok-to-stop since it will always fail and eventually force
pass but let's not wait for nothing.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-25 09:30:38 +02:00
Travis Nielsen cdfe1982d1 ceph: consolidate the calls to set mon config
The mon config had two different implementations that have evolved
over the lifetime of the project. This is a simple refactor to remove
the SetConfig() option and stick with the MonStore as a single
implementation for updating the mon store.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-08-24 14:38:26 -06:00
parth-gr 1601ba1c1c core: convert util.NewSet() to sets.NewString()
Converting util.NewSet() instance to use sets.NewString() instance

Closes: https://github.com/rook/rook/issues/8479
Signed-off-by: parth-gr <paarora@redhat.com>
2021-08-20 18:56:25 +05:30
Sébastien Han bb4e572d2f Merge pull request #8465 from parth-gr/import-rbd-mirror
ceph: use mktemp file in ImportRBDMirrorBootstrapPeer
2021-08-05 09:44:43 +02:00
parth-gr 48918c7cf0 ceph: use mktemp file in ImportRBDMirrorBootstrapPeer
Do not point to a made-up file in /tmp
Use ioutil.TempFile() for creating a token file

Closes: https://github.com/rook/rook/issues/8446
Signed-off-by: parth-gr <paarora@redhat.com>
2021-08-04 19:22:34 +05:30
Sébastien Han 8fdb9fe2b9 ceph: remove old network types
Those types are legacy, not really used anywhere and not needed anymore.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-04 10:56:01 +02:00
Sébastien Han 412b3eaf5e Merge pull request #8447 from leseb/fix-8438
ceph: ignore errors when mirroring is not enabled on the filesystem
2021-08-03 16:06:16 +02:00
Sébastien Han 7318423506 ceph: ignore errors when mirroring is not enabled on the filesystem
Trying to disable mirroring on a cluster where mirroring is not enabled
results in an error during upgrades. Let's catch this error and ignore
it.

Funny enough the exec error resembles to:

```
Error ENOTSUP: Module 'mirroring' is not enabled (required by command 'fs snapshot mirror disable'): use `ceph mgr module enable mirroring` to enable it: exit status 95
```

So we get ENOTSUP which is 45 but exit status 95...

Closes: https://github.com/rook/rook/issues/8438
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-03 15:33:04 +02:00
Sébastien Han 8556f3fe26 ceph: remove pool id from the peer
We don't need to put this information in the token.
It's not useful and not used anywhere.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-03 09:44:54 +02:00
Sébastien Han 630c2f6a8b ceph: add an rbd-mirror bootstrap token on cluster creation
They are scenarios where the mirroring information want to be shared
between clusters prior to creating pool. Because the bootstrap peer
import command needs a pool name to operate this is not suitable. So
additionally now each time the cluster is reconciled and on any new
clusters a new secret will be created that contains a boostrap peer
token. It can be exchanged with another cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-02 18:34:27 +02:00
Travis Nielsen f5030f0560 ceph: remove old references to mimic version
References to luminous and mimic, and unnecessary references to nautilus
were found in the docs that have been cleaned up and clarified. At the same
time, the tests are updated to use more recent versions.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-28 14:00:35 -06:00
Sébastien Han a62e7e75e7 Merge pull request #8380 from sp98/update-cephblockpool-cr
ceph: add mirroring peer specification to CephBlockPool CRD
2021-07-28 17:34:52 +02:00
Santosh Pillai 3b50f5fd00 ceph: add MirroringPeerSpec to CephBlockPool.Spec.Mirroring
Instead of having the RBDMirroringPeerSpec in the RBDMirror CRD we want
to move it on the CephBlockPool CRD so that each pool can have its own peer.
This enables pool to re-use an existing peer secret if it points to the
same cluster peer.So we don't need to duplicate secret peer anymore.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-07-28 17:47:36 +05:30
Satoru Takeuchi 30e4fbb01f ceph: make the timeout of ceph commands cofigurable
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-27 12:42:16 +00:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Sébastien Han 7c5a86cb85 ceph: remove operator config ROOK_ALLOW_MULTIPLE_FILESYSTEMS
Multiple filesystems are supported as of the Ceph Pacific release. So we
remove the flag from the operator configuration and just have a ceph
version check instead.

The Operator configuration option `ROOK_ALLOW_MULTIPLE_FILESYSTEMS`
has been removed in favor of simply verifying the Ceph version is
at least Pacific.
Multiple filesystems are stable since Ceph Pacific.
So users who had `ROOK_ALLOW_MULTIPLE_FILESYSTEMS` enabled will
need to update their Ceph version to Pacific.

Closes: https://github.com/rook/rook/issues/7183
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-22 11:42:33 +02:00
Travis Nielsen 14227600cc Merge pull request #7535 from travisn/mon-stretch-location
ceph: Set the location on the mon daemon for stretch clusters
2021-07-20 08:17:01 -06:00
Sébastien Han 1900160939 ceph: proxy rbd commands when multus is enabled
We now pass the networking spec to the clusterInfo so that the executor
can make the right decision on how to execute a command.
This is a small change that allows us to remove the CephBlockPool CR
since it's checking for rbd images. On a Multus deployment, the operator
does not have the network annotations, thus has no access to the OSD
network then rbd commands are hanging forever.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-20 12:28:00 +02:00
Travis Nielsen fca761c950 ceph: set the location on the mon daemon
The mon daemon in a stretch cluster now can have its location set
as a CLI param instead of setting it with a separate command.
This enables mon failover to set the location of a mon immediately
when it is joining quorum instead of having a delayed command
to set the location.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-19 14:30:43 -06:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Sébastien Han baaea4a1ea Merge pull request #7604 from leseb/cephfs-mirror-peer-config
ceph: add filesystem mirror peers configuration
2021-07-05 10:36:30 +02:00
Santosh Pillai 05b0ae09bb ceph: disable mirroring
The PR disables mirroring on a pool when PoolSpec.Mirroring.Enabled
is set to false
- disable mirroring on the pool for `Mirroring.Mode == pool`
- Add warning for the user to disable mirroring manually for `Mirroring.Mode == image`
- Stop the mirroring health checker goroutine
- Reset the mirroring health check status in the CR.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-07-02 10:35:38 +05:30
Sébastien Han b578f916e7 ceph: add fs mirror config
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.

So the automatic configuration of Ceph Filesystem peers is now possible.

By editing the CephFilesystem CRD, you can now turn on mirroring:

```yaml
  mirroring:
    enabled: false
    # list of Kubernetes Secrets containing the peer token
    # for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
    peers:
      secretNames:
        - secondary-cluster-peer
```

Also, the mirroring status is displayed in the CR status:

```
status:
  info:
    fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
  mirroringStatus:
    daemonsStatus:
    - daemon_id: 4186
      filesystems:
      - filesystem_id: 2
        name: myfs
    lastChecked: "2021-07-01T14:16:29Z"
  phase: Ready
  snapshotScheduleStatus:
    lastChecked: "2021-07-01T14:16:29Z"
    snapshotSchedules:
    - fs: myfs
      path: /
      rel_path: /
      retention: {}
      schedule: 24h
```

Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 17:35:19 +02:00
Santosh Pillai 4041868e2e ceph: hybrid Storage Pools
Creates hybrid crush rule for choosing Primary OSD for
high performing SSD devices and remaining OSD for low performance HDD devices.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-07-01 11:16:54 +05:30
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Andy Bursavich 11f0add706 build: update golangci-lint version and nolint comment format
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:50:04 -07:00
Travis Nielsen c961d9f7ac Merge pull request #8117 from degorenko/device-class-for-present-osds
ceph: set device class for already present osd deployments
2021-06-16 12:07:11 -06:00
subhamkrai 5d5bdbbbda ceph: remove --force when creating filesystem
we should not use --force when creating filesystem(fs)
as it will recreate fs which will leads to data loss
when `preserveFilesystemOnDelete` is set to 'true'.

this commit removes --force argument when create fs.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-06-16 13:39:49 +05:30