Commit Graph
245 Commits
Author SHA1 Message Date
Sébastien Han 7402c2cce6 osd: check if osd is ok-to-stop before removal
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.

So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-25 10:21:32 +01:00
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Jiffin Tony Thottan aba50d3ca9 object: add support in RGW to communicate vault with TLS
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-11-17 10:19:28 +05:30
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
Yuichiro Ueno 0b575703c7 core: add context parameter to k8sutil deployment
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 14:58:13 +09:00
Travis Nielsen 427996a7c0 osd: increase wait timeout for osd prepare cleanup
When a reconcile is started for OSDs, the prepare jobs are first
deleted from a previous reconcile. The timeout for the osd prepare
job deletion was only 40s. After that timeout, the reconcile attempts
to continue waiting for the pod, but of course will never complete
since the OSD prepare was not running in the first place, causing the
reconcile to wait indefinitely. In the reported issue, the osd prepare
jobs were actually deleted successfully, the timeout just wasn't long
enough. Pods need at least a minute to be forcefully deleted,
so we increase the timeout to 90s to give it some extra buffer.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-05 14:20:43 -06:00
Sébastien Han d06a6f93a5 core: run operator with rook user
The rook operator as well as the toolbox pod run with the "rook" user
with UID 2016. The UID was chosen based on the year of the initial
commit in the rook/rook repository.
No more root user running.

Closes: https://github.com/rook/rook/issues/8734
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-05 17:22:04 +01:00
Sébastien Han 16729e0e38 osd: use multiple service account for vault role
We can pass bound_service_account_names with a comma separated list of
service accounts. Let's do this instead of remapping new values.
Earlier, we thought a single service account could be added per Vault
role and we were using other variables like
`VAULT_AUTH_KUBERNETES_ROOK_OPERATOR_ROLE` that we were remapping to
`VAULT_AUTH_KUBERNETES_ROLE` internal for the API calls to Vault.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-04 15:15:46 +01:00
Sébastien Han 3b5bbf28d5 Merge pull request #9003 from leseb/vault-k8s-auth
osd: add support for k8s with vault kms
2021-10-25 18:02:48 +02:00
Sébastien Han 18a4047679 osd: add support for k8s with vault kms
Rook cluster-wide encryption can now use the native Kubernetes
authentication to interact with vault KMS instead of using the token
method.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-21 13:59:28 +02:00
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
Sébastien Han 4017f94464 osd: use rook binary to fetch key encryption key
When deploying a cluster-wde encrypted cluster we now use the rook
binary to execute some code to fetch the key encryption key.

Using cURL all the time to fetch the key has its limitations. The
incoming integration with Kubernetes Authentication through service
accounts is leading the usage of cURL to its end.
The logic is really complex and prone to errors. Re-implementing the
Go logic into Bash is not a viable option.

Also, using the lib in a binary allows us to keep a consistent behavior
throughout the life cycle of our code.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-19 17:29:25 +02:00
Sébastien Han 61afadd97e ceph: fix kms auto-detection when full TLS
When TLS is used and includes a caert, client key/cert, we need to copy
the content of the secret to a file in the operator's container
filesystem so that we can build the TLS config and thus the HTTP Client,
which reads those files.
Also, removing the files after each API call so they don't persist on
the filesystem forever.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-13 18:12:58 +02:00
Sébastien Han 2e73baf4f5 ceph: do not fail on keys deletion
Prior, we were returning a nil map and thus the assignment for forced
deletion was not working since we were trying to assign on a nil map.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-29 15:59:18 +02:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Jonas Zeiger 0e72a7c2bf ceph: fix lvm osd db device check
Signed-off-by: Jonas Zeiger <jonas.zeiger@talpidae.net>
2021-09-13 12:19:16 +02:00
Hiroya Onoe 2ff5413b75 ceph: fix error message in UpdateNodeStatus
The error message in UpdateNodeStatus regards the second argument
as node name. However, it is a PVC name in OSD on PVC.

Signed-off-by: Hiroya Onoe <onoehiroya@gmail.com>
2021-09-02 02:43:16 +00:00
Sébastien Han d675969567 ceph: fix vault kv secret engine auto-detection
Passing struct by value essentially gives you a copy, so when modified
within a function, the scope is then reduced to that function. Using
pointers solves that you mutate the struct as many times as you want from
anywhere.
As a result, the auto-detection of the Vault KV backend was not working
correctly.
Also, added a ton of unit tests for Vault.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-31 15:29:04 +02:00
Satoru Takeuchi 2218c53fda Merge pull request #8403 from cybozu-go/ceph-dont-remove-pvc-in-osd-purge-job
ceph: add an option to preserve pvc in osd purge job
2021-07-30 01:04:28 +09:00
Satoru Takeuchi 89b6a6028c ceph: add an option to preserve pvc in osd purge job
Sometimes we want to investigate a PVC for removed OSD.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-29 14:20:06 +00:00
Sébastien Han 99e00dea1e ceph: auto detect vault k/v version
Rook will now auto detect the kv version of the vault server. This
allows users not having to pass the VAULT_BACKEND configuration in the
CephCluster CR.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-29 10:16:31 +02:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Blaine Gardner c8a2db368a ceph: disable raw mode for disks
Even if we can use raw mode, do NOT use raw mode on disks. Ceph bluestore disks can
sometimes appear as though they have "phantom" Atari (AHDI) partitions created on them
when they don't in reality. This is due to a series of bugs in the Linux kernel when it
is built with Atari support enabled. This behavior does not appear for raw mode OSDs on
partitions, and we need the raw mode to create partition-based OSDs. We cannot merely
skip creating OSDs on "phantom" partitions due to a bug in `ceph-volume raw inventory`
which reports only the phantom partitions (and malformed OSD info) when they exist and
ignores the original (correct) OSDs created on the raw disk.

Resolves https://github.com/rook/rook/issues/7940

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-07-22 09:59:56 -06:00
Sébastien Han 7e0b58cc27 ceph: destroy all vault keys on kv version 2
By default, keys stored in Vault KV Secret Engine are versioned and keys
are soft-deleted, see Vault documentation:

When deleting data the standard vault kv delete command will perform a soft
delete.
It will mark the version as deleted and populate a deletion_time timestamp.
Soft deletes do not remove the underlying version data from storage,
which allows the version to be undeleted.
The vault kv undelete command handles undeleting versions. A version's
data is permanently deleted only when the key has more versions than are
allowed by the max-versions setting, or when using vault kv destroy.
When the destroy command is used the underlying version data will be
removed and the key metadata will be marked as destroyed.
If a version is cleaned up by going over max-versions the version
metadata will also be removed from the key.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-09 11:03:17 +02:00
Blaine Gardner 7f516b9e3d ceph: ignore atari partitions when scanning disks
Ceph bluestore raw disks can sometimes appear as though they have Atari
(AHDI) partitions on them. If a disk has Atari partitions, we just
ignore them as though they don't exist. This should be a safe assumption
since the hardware was last manufactured in 1992 and likely can't run
Kubernetes.

If we don't ignore the Atari partitions, Rook can create a new OSD on a
disk that is already running an OSD, corrupting the first and possibly
also the latest OSD. This can happen an arbitrary number of times per
disk.

More info: https://github.com/rook/rook/issues/7940

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-30 09:42:33 -06:00
Travis Nielsen 06fbf29b3c Merge pull request #8028 from degorenko/device-classes-resource-limits
ceph: add ability to set resource limit for OSDs based on device classes
2021-06-08 09:01:14 -06:00
Denis Egorenko 8a556fd847 ceph: add ability to set resource limit for OSDs based on device classes
Currently it is not possible to set resource limits based on device classes
for different OSDs. Now adding an ability to use predefined keys for main
cluster spec Resource section to reflect resource limits for different OSDs
with different device classes.

Closes: https://github.com/rook/rook/issues/8007
Signed-off-by: Denis Egorenko <degorenko@mirantis.com>
2021-06-08 17:46:41 +04:00
Sébastien Han e38601ac2d ceph: always skip disks with filesystem
Unless they are encrypted disk and have markers for a ceph cluster.

Closes: https://github.com/rook/rook/issues/8046
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-03 14:42:29 +02:00
Sébastien Han 26c0f8bcc1 ceph: only pass vault volume projection is necessary
We don't need to pass the Volume with projection for TLS when TLS is not enabled
Somehow when this happens and we try to update a deployment spec it fails with:

```
 ValidationError(Pod.spec.volumes[7].projected): missing required field "sources"
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-31 16:33:40 +02:00
Sébastien Han 0004869f66 ceph: add ceph cluster fsid to LUKS header
When configuring encrypted on-pvc clusters we now set the cluster fsid
in the LUKS header. If needed we can then determine if the encrypted
block is part of our cluster or not. We also attach the pvc_name in case
it might become useful.

Closes: https://github.com/rook/rook/issues/7991
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-31 15:01:52 +02:00
Travis Nielsen 55dc1d2ba5 ceph: fix intermittent volume unit test
The unit test TestConfigureCVDevices was failing intermittently
due to comparing the first element of the slice of two elements.
The slice ordering is not guaranteed, which caused a failure
intermittently. The test only compares the critical values in the test
and no need to check elements of the slice.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-05 19:12:28 -06:00
Travis Nielsen a53a270923 Merge pull request #7823 from travisn/filestore-refs
ceph: Remove obsolete references to filestore
2021-05-05 09:10:43 -06:00
Travis Nielsen d5374f11a1 ceph: set the device class on raw mode osds
The deviceClass was only being set for LVM mode OSDs and was
missed for the raw mode for non-PVC. Now the deviceClass will be set
for all expected scenarios.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-04 17:13:49 -06:00
Travis Nielsen eba91a4bd6 ceph: remove obsolete references to filestore
Filestore support has been long gone with only bluestore
currently supported.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-04 15:50:20 -06:00
Jiffin Tony Thottan 8e5c6c126c ceph: modify ValidateConnectionDetails() in kms
The ValidateConnectionDetails() contains `clusterSpec` assuming `SecuritySpec` only
part to `cephCluster` CRD. But we may need to validate same for `SecuritySpec` in
`cephObjectStore` CRD. Hence modifying the signature of the api.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-28 13:03:16 +05:30
Sébastien Han 1930eb3b4c ceph: print provision error entirely
If the ceph-volume prepare call fails, let's print the output so we can
see it in the operator's logs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-04-22 11:30:37 +02:00
Blaine Gardner b4946aba95 ceph: use ceph-volume v2 report format for Pacific
This fixes cases where raw provisioning mode still cannot be used for
Rook v1.6 and Ceph Pacific clusters where json.Unmarshal fails to read
the ceph-volume output after OSD provisioning.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-04-20 12:45:05 -06:00
Shachar Sharon 9766e9b8fb ceph: allow passing 'osd-crush-initial-weight'
Ceph support the option '--osd-crush-initial-weight' upon OSD start,
which sets an explicit weight (in TiB units) to specific OSD. Allow
passing this option all the way from the user (similar to
'DeviceClass'), for the special case where end users wants it cluster
to have non-even balance over specific OSDs (e.g., one of the OSDs is
placed over a partition alongside OS-partition).

ROOK issue: https://github.com/rook/rook/issues/7448

Signed-off-by: Shachar Sharon <ssharon@redhat.com>
2021-04-20 20:50:54 +03:00
Lars Lehtonen bb323de403 ceph: fix multiple imports
This mops up the last of the multiple-imports within ceph-related
packages, in pkg/daemon/ceph/agent/flexvolume/manager/ceph and
pkg/daemon/ceph/osd.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-19 00:06:10 -07:00
Satoru Takeuchi a4fb080590 ceph: continue to get available devices if failed to get a device info
getAvailableDevices() should continue if something wrong happens in a device than returning
immediately with error.

Closes: https://github.com/rook/rook/issues/7543

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-04-14 03:13:06 +00:00
Blaine Gardner 36e98aa44d Revert "ceph: remove auth in osd-purge job"
This reverts commit 80d2cde5c7.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-04-07 13:12:19 -06:00
Blaine Gardner 80d2cde5c7 ceph: remove auth in osd-purge job
`ceph osd purge` should remove the OSD auth, but many users have issues
where they are unable to create new OSDs after removing old ones due to
the old OSD auth still being present. Therefore, run `ceph auth del` for
the OSD after purging to fix this.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-04-06 16:59:33 -06:00
Blaine Gardner 1df4336a68 Merge pull request #7386 from BlaineEXE/update-osds-in-parallel
ceph: Update osds in parallel
2021-03-29 16:32:07 -06:00
Blaine Gardner 795124b7a8 ceph: update osds in parallel
Update OSDs in parallel per the design in
design/ceph/update-osds-in-parallel.md

The max number of OSDs updated in parallel is currently fixed at 20.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-03-29 10:55:28 -06:00
Sébastien Han 29c55575b7 ceph: avoid restarting all encrypted osd on cluster growth
When the cluster is growing, the operator always processes all OSDs and
re-generate their Deployment spec. The generated environment variables
for encryption are coming from a map, which by design is unordered. So
once generated the deployment will see their spec change since the
environment variable position will change.

See from the following diff:

```
        env:							        env:
        - name: KMS_SERVICE_NAME			      <
          value: vault					      <
        - name: VAULT_NAMESPACE				      <
          value: ocsns					      <
        - name: VAULT_TLS_SERVER_NAME			      <
        - name: VAULT_BACKEND_PATH			      <
          value: n_ocs					      <
        - name: VAULT_ADDR					        - name: VAULT_ADDR
          value: https://vault.qe.rh-ocs.com:8200		          value: https://vault.qe.rh-ocs.com:8200
							      >	        - name: VAULT_BACKEND_PATH
							      >	          value: n_ocs
							      >	        - name: VAULT_TLS_SERVER_NAME
        - name: KMS_PROVIDER					        - name: KMS_PROVIDER
          value: vault						          value: vault
							      >	        - name: KMS_SERVICE_NAME
							      >	          value: vault
							      >	        - name: VAULT_NAMESPACE
							      >	          value: ocsns
        - name: VAULT_TOKEN					        - name: VAULT_TOKEN
          valueFrom:						          valueFrom:
            secretKeyRef:					            secretKeyRef:
              key: token					              key: token
              name: ocs-kms-token				              name: ocs-kms-token
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-26 16:16:13 +01:00
Sébastien Han dd275a2906 ceph: set default vault backend path when not specified
If the path was not specified, the OSD init container would crash loop,
failing to reach the vault path.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-22 16:56:41 +01:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Travis Nielsen c0eacc89d5 Merge pull request #7235 from travisn/mgr-sidecar
ceph: Allow two mgr daemons and actively reconcile the mgr services with a sidecar
2021-03-10 10:35:40 -07:00
Travis Nielsen 388ff3e78b ceph: start sidecar to monitor active mgr
The mgr daemon may be failed over by ceph if the active mgr is not
responding and the standby mgr is available. If the active mgr changes
the services for the dashboard and metrics will be updated with a
label selector for the new active mgr. The services cannot direct
traffic to the standby mgr or else they will be incorrectly redirected
to the active mgr.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:37:17 -07:00
Sébastien Han a51d8c5f2e ceph: use raw mode for pacific on non-pvc
We will now use raw mode on non-pvc for pacific and onward.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-10 12:18:31 +01:00