Commit Graph
585 Commits
Author SHA1 Message Date
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Jonas Zeiger 0e72a7c2bf ceph: fix lvm osd db device check
Signed-off-by: Jonas Zeiger <jonas.zeiger@talpidae.net>
2021-09-13 12:19:16 +02:00
Hiroya Onoe 2ff5413b75 ceph: fix error message in UpdateNodeStatus
The error message in UpdateNodeStatus regards the second argument
as node name. However, it is a PVC name in OSD on PVC.

Signed-off-by: Hiroya Onoe <onoehiroya@gmail.com>
2021-09-02 02:43:16 +00:00
Sébastien Han d4dd0577f9 Merge pull request #8529 from sp98/id-mapping
ceph: add ClusterID and PoolID mappings between local and peer cluster
2021-08-31 17:32:14 +02:00
Santosh Pillai 3f8abec403 ceph: add ClusterID and PoolID mappings between local and peer cluster
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters

This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-08-31 20:08:56 +05:30
Sébastien Han d675969567 ceph: fix vault kv secret engine auto-detection
Passing struct by value essentially gives you a copy, so when modified
within a function, the scope is then reduced to that function. Using
pointers solves that you mutate the struct as many times as you want from
anywhere.
As a result, the auto-detection of the Vault KV backend was not working
correctly.
Also, added a ton of unit tests for Vault.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-31 15:29:04 +02:00
Sébastien Han 9744f0f2c6 Merge pull request #8583 from leseb/cancel-retry
ceph: do not check ok-to-stop when OSDs are in CLBO
2021-08-25 17:38:06 +02:00
Sébastien Han 88048cc4e8 ceph: do not check ok-to-stop when OSDs are in CLBO
This handles the scenario where the OSDs have been created but not yet
started due to a wrong CR configuration.
For instance, when OSDs are encrypted and Vault is used to store
encryption keys, if the KV version is incorrect during the cluster
initialization the OSDs will fail to start and stay in CLBO until the
CR is updated again with the correct KV version so that it can start.

For this scenario, if the CRUSH map has no host registered yet it's fair
to assume the initialization broke and we need to fix it. So when don't
need to call ok-to-stop since it will always fail and eventually force
pass but let's not wait for nothing.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-25 09:30:38 +02:00
Travis Nielsen cdfe1982d1 ceph: consolidate the calls to set mon config
The mon config had two different implementations that have evolved
over the lifetime of the project. This is a simple refactor to remove
the SetConfig() option and stick with the MonStore as a single
implementation for updating the mon store.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-08-24 14:38:26 -06:00
parth-gr 1601ba1c1c core: convert util.NewSet() to sets.NewString()
Converting util.NewSet() instance to use sets.NewString() instance

Closes: https://github.com/rook/rook/issues/8479
Signed-off-by: parth-gr <paarora@redhat.com>
2021-08-20 18:56:25 +05:30
Sébastien Han bb4e572d2f Merge pull request #8465 from parth-gr/import-rbd-mirror
ceph: use mktemp file in ImportRBDMirrorBootstrapPeer
2021-08-05 09:44:43 +02:00
parth-gr 48918c7cf0 ceph: use mktemp file in ImportRBDMirrorBootstrapPeer
Do not point to a made-up file in /tmp
Use ioutil.TempFile() for creating a token file

Closes: https://github.com/rook/rook/issues/8446
Signed-off-by: parth-gr <paarora@redhat.com>
2021-08-04 19:22:34 +05:30
Sébastien Han 8fdb9fe2b9 ceph: remove old network types
Those types are legacy, not really used anywhere and not needed anymore.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-04 10:56:01 +02:00
Sébastien Han 412b3eaf5e Merge pull request #8447 from leseb/fix-8438
ceph: ignore errors when mirroring is not enabled on the filesystem
2021-08-03 16:06:16 +02:00
Sébastien Han 7318423506 ceph: ignore errors when mirroring is not enabled on the filesystem
Trying to disable mirroring on a cluster where mirroring is not enabled
results in an error during upgrades. Let's catch this error and ignore
it.

Funny enough the exec error resembles to:

```
Error ENOTSUP: Module 'mirroring' is not enabled (required by command 'fs snapshot mirror disable'): use `ceph mgr module enable mirroring` to enable it: exit status 95
```

So we get ENOTSUP which is 45 but exit status 95...

Closes: https://github.com/rook/rook/issues/8438
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-03 15:33:04 +02:00
Sébastien Han 8556f3fe26 ceph: remove pool id from the peer
We don't need to put this information in the token.
It's not useful and not used anywhere.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-03 09:44:54 +02:00
Sébastien Han 630c2f6a8b ceph: add an rbd-mirror bootstrap token on cluster creation
They are scenarios where the mirroring information want to be shared
between clusters prior to creating pool. Because the bootstrap peer
import command needs a pool name to operate this is not suitable. So
additionally now each time the cluster is reconciled and on any new
clusters a new secret will be created that contains a boostrap peer
token. It can be exchanged with another cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-02 18:34:27 +02:00
Satoru Takeuchi 2218c53fda Merge pull request #8403 from cybozu-go/ceph-dont-remove-pvc-in-osd-purge-job
ceph: add an option to preserve pvc in osd purge job
2021-07-30 01:04:28 +09:00
Satoru Takeuchi 89b6a6028c ceph: add an option to preserve pvc in osd purge job
Sometimes we want to investigate a PVC for removed OSD.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-29 14:20:06 +00:00
Travis Nielsen 990d92790b Merge pull request #8420 from travisn/cleanup-old-versions
ceph: Remove old references to luminous and mimic versions
2021-07-29 07:40:46 -06:00
Sébastien Han 99e00dea1e ceph: auto detect vault k/v version
Rook will now auto detect the kv version of the vault server. This
allows users not having to pass the VAULT_BACKEND configuration in the
CephCluster CR.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-29 10:16:31 +02:00
Travis Nielsen f5030f0560 ceph: remove old references to mimic version
References to luminous and mimic, and unnecessary references to nautilus
were found in the docs that have been cleaned up and clarified. At the same
time, the tests are updated to use more recent versions.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-28 14:00:35 -06:00
Sébastien Han a62e7e75e7 Merge pull request #8380 from sp98/update-cephblockpool-cr
ceph: add mirroring peer specification to CephBlockPool CRD
2021-07-28 17:34:52 +02:00
Santosh Pillai 3b50f5fd00 ceph: add MirroringPeerSpec to CephBlockPool.Spec.Mirroring
Instead of having the RBDMirroringPeerSpec in the RBDMirror CRD we want
to move it on the CephBlockPool CRD so that each pool can have its own peer.
This enables pool to re-use an existing peer secret if it points to the
same cluster peer.So we don't need to duplicate secret peer anymore.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-07-28 17:47:36 +05:30
Satoru Takeuchi 30e4fbb01f ceph: make the timeout of ceph commands cofigurable
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-27 12:42:16 +00:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Travis Nielsen 16403ac901 Merge pull request #8350 from leseb/fix-7183
ceph: remove operator config `ROOK_ALLOW_MULTIPLE_FILESYSTEMS`
2021-07-22 11:33:40 -06:00
Travis Nielsen fd141af92c Merge pull request #8319 from BlaineEXE/disable-raw-mode-for-disks
ceph: disable raw mode for disks
2021-07-22 11:17:54 -06:00
Blaine Gardner c8a2db368a ceph: disable raw mode for disks
Even if we can use raw mode, do NOT use raw mode on disks. Ceph bluestore disks can
sometimes appear as though they have "phantom" Atari (AHDI) partitions created on them
when they don't in reality. This is due to a series of bugs in the Linux kernel when it
is built with Atari support enabled. This behavior does not appear for raw mode OSDs on
partitions, and we need the raw mode to create partition-based OSDs. We cannot merely
skip creating OSDs on "phantom" partitions due to a bug in `ceph-volume raw inventory`
which reports only the phantom partitions (and malformed OSD info) when they exist and
ignores the original (correct) OSDs created on the raw disk.

Resolves https://github.com/rook/rook/issues/7940

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-07-22 09:59:56 -06:00
Sébastien Han 7c5a86cb85 ceph: remove operator config ROOK_ALLOW_MULTIPLE_FILESYSTEMS
Multiple filesystems are supported as of the Ceph Pacific release. So we
remove the flag from the operator configuration and just have a ceph
version check instead.

The Operator configuration option `ROOK_ALLOW_MULTIPLE_FILESYSTEMS`
has been removed in favor of simply verifying the Ceph version is
at least Pacific.
Multiple filesystems are stable since Ceph Pacific.
So users who had `ROOK_ALLOW_MULTIPLE_FILESYSTEMS` enabled will
need to update their Ceph version to Pacific.

Closes: https://github.com/rook/rook/issues/7183
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-22 11:42:33 +02:00
Travis Nielsen 14227600cc Merge pull request #7535 from travisn/mon-stretch-location
ceph: Set the location on the mon daemon for stretch clusters
2021-07-20 08:17:01 -06:00
Sébastien Han 1900160939 ceph: proxy rbd commands when multus is enabled
We now pass the networking spec to the clusterInfo so that the executor
can make the right decision on how to execute a command.
This is a small change that allows us to remove the CephBlockPool CR
since it's checking for rbd images. On a Multus deployment, the operator
does not have the network annotations, thus has no access to the OSD
network then rbd commands are hanging forever.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-20 12:28:00 +02:00
Travis Nielsen fca761c950 ceph: set the location on the mon daemon
The mon daemon in a stretch cluster now can have its location set
as a CLI param instead of setting it with a separate command.
This enables mon failover to set the location of a mon immediately
when it is joining quorum instead of having a delayed command
to set the location.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-19 14:30:43 -06:00
Sébastien Han 7e0b58cc27 ceph: destroy all vault keys on kv version 2
By default, keys stored in Vault KV Secret Engine are versioned and keys
are soft-deleted, see Vault documentation:

When deleting data the standard vault kv delete command will perform a soft
delete.
It will mark the version as deleted and populate a deletion_time timestamp.
Soft deletes do not remove the underlying version data from storage,
which allows the version to be undeleted.
The vault kv undelete command handles undeleting versions. A version's
data is permanently deleted only when the key has more versions than are
allowed by the max-versions setting, or when using vault kv destroy.
When the destroy command is used the underlying version data will be
removed and the key metadata will be marked as destroyed.
If a version is cleaned up by going over max-versions the version
metadata will also be removed from the key.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-09 11:03:17 +02:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Sébastien Han baaea4a1ea Merge pull request #7604 from leseb/cephfs-mirror-peer-config
ceph: add filesystem mirror peers configuration
2021-07-05 10:36:30 +02:00
Santosh Pillai 05b0ae09bb ceph: disable mirroring
The PR disables mirroring on a pool when PoolSpec.Mirroring.Enabled
is set to false
- disable mirroring on the pool for `Mirroring.Mode == pool`
- Add warning for the user to disable mirroring manually for `Mirroring.Mode == image`
- Stop the mirroring health checker goroutine
- Reset the mirroring health check status in the CR.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-07-02 10:35:38 +05:30
Sébastien Han b578f916e7 ceph: add fs mirror config
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.

So the automatic configuration of Ceph Filesystem peers is now possible.

By editing the CephFilesystem CRD, you can now turn on mirroring:

```yaml
  mirroring:
    enabled: false
    # list of Kubernetes Secrets containing the peer token
    # for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
    peers:
      secretNames:
        - secondary-cluster-peer
```

Also, the mirroring status is displayed in the CR status:

```
status:
  info:
    fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
  mirroringStatus:
    daemonsStatus:
    - daemon_id: 4186
      filesystems:
      - filesystem_id: 2
        name: myfs
    lastChecked: "2021-07-01T14:16:29Z"
  phase: Ready
  snapshotScheduleStatus:
    lastChecked: "2021-07-01T14:16:29Z"
    snapshotSchedules:
    - fs: myfs
      path: /
      rel_path: /
      retention: {}
      schedule: 24h
```

Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 17:35:19 +02:00
Santosh Pillai 4041868e2e ceph: hybrid Storage Pools
Creates hybrid crush rule for choosing Primary OSD for
high performing SSD devices and remaining OSD for low performance HDD devices.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-07-01 11:16:54 +05:30
Blaine Gardner 037945aabc Merge pull request #8112 from BlaineEXE/dependencies-2-son-of-dependencies
ceph: block delete object store when buckets and/or users exist
2021-06-30 12:45:58 -06:00
Blaine Gardner 7f516b9e3d ceph: ignore atari partitions when scanning disks
Ceph bluestore raw disks can sometimes appear as though they have Atari
(AHDI) partitions on them. If a disk has Atari partitions, we just
ignore them as though they don't exist. This should be a safe assumption
since the hardware was last manufactured in 1992 and likely can't run
Kubernetes.

If we don't ignore the Atari partitions, Rook can create a new OSD on a
disk that is already running an OSD, corrupting the first and possibly
also the latest OSD. This can happen an arbitrary number of times per
disk.

More info: https://github.com/rook/rook/issues/7940

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-30 09:42:33 -06:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Andy Bursavich 11f0add706 build: update golangci-lint version and nolint comment format
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:50:04 -07:00
Andy Bursavich b801ff7da0 ceph: extract flexvolume dependency on k8s.io/kubernetes
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:36:08 -07:00
Travis Nielsen c961d9f7ac Merge pull request #8117 from degorenko/device-class-for-present-osds
ceph: set device class for already present osd deployments
2021-06-16 12:07:11 -06:00
subhamkrai 5d5bdbbbda ceph: remove --force when creating filesystem
we should not use --force when creating filesystem(fs)
as it will recreate fs which will leads to data loss
when `preserveFilesystemOnDelete` is set to 'true'.

this commit removes --force argument when create fs.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-06-16 13:39:49 +05:30
Denis Egorenko 6dbb902674 ceph: set device class for already present osd deployments
Add ability to set device class for existing osd deployments,
based on its crush map device class.

Related-Issue: rook#8028
Signed-off-by: Denis Egorenko <degorenko@mirantis.com>
2021-06-15 13:31:37 +04:00
Travis Nielsen b1f4411fb4 ceph: disable insecure global id if no insecure clients
In the latest Ceph releases starting with v16.2.1, all clients are recommended
to be updated so they will have a security fix to connect with a secure
global ID. A health warning will be raised if any insecure clients are connected
and another health warning is raised if insecure clients are still allowed.
Rook will now disable allowing the insecure clients if the health warning
is not being raised to indicate that there are insecure clients still connected.
This means that upgraded clusters will not have this disabled until all the
daemons are updated.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-06-10 10:05:30 -06:00
Travis Nielsen 06fbf29b3c Merge pull request #8028 from degorenko/device-classes-resource-limits
ceph: add ability to set resource limit for OSDs based on device classes
2021-06-08 09:01:14 -06:00
Denis Egorenko 8a556fd847 ceph: add ability to set resource limit for OSDs based on device classes
Currently it is not possible to set resource limits based on device classes
for different OSDs. Now adding an ability to use predefined keys for main
cluster spec Resource section to reflect resource limits for different OSDs
with different device classes.

Closes: https://github.com/rook/rook/issues/8007
Signed-off-by: Denis Egorenko <degorenko@mirantis.com>
2021-06-08 17:46:41 +04:00