The failover of the arbiter mon in a stretch cluster was sometimes
failing due to the new tiebreaker not being set in ceph.
Rook would repeatedly try to remove the old tiebreaker mon
and keep failing because the new tiebreaker had not been set.
Now we make setting the tiebreaker idempotent in case the operator
restarts in the middle of the operation or some other corner
case causes the expected tiebreaker to be set. In that case,
the next reconcile will also ensure the tiebreaker mon is
set as expected.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Sometimes Ceph uses a different standard output to return errors or
merges standard error to standard out. So let's allow some commands to
return both in the output.
Signed-off-by: Sébastien Han <seb@redhat.com>
Some users have reported issues while adding the token, this is not
always reproducable so perhaps it's a typo when importing the token and
adding trailing spaces.
Closes: https://github.com/rook/rook/issues/9151
Signed-off-by: Sébastien Han <seb@redhat.com>
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.
So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.
Signed-off-by: Sébastien Han <seb@redhat.com>
For CRD not using the new nfs spec that includes the pool settings,
applying the "size" property won't work since it is set to 0. The pool
still gets created but returns an error. The loop is re-queued but on
the second run the pool is detected so no further configuration is done.
Closes: https://github.com/rook/rook/issues/9205
Signed-off-by: Sébastien Han <seb@redhat.com>
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil pod functions. By this, we
can handle cancellation during API call of pod resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
When a reconcile is started for OSDs, the prepare jobs are first
deleted from a previous reconcile. The timeout for the osd prepare
job deletion was only 40s. After that timeout, the reconcile attempts
to continue waiting for the pod, but of course will never complete
since the OSD prepare was not running in the first place, causing the
reconcile to wait indefinitely. In the reported issue, the osd prepare
jobs were actually deleted successfully, the timeout just wasn't long
enough. Pods need at least a minute to be forcefully deleted,
so we increase the timeout to 90s to give it some extra buffer.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The rook operator as well as the toolbox pod run with the "rook" user
with UID 2016. The UID was chosen based on the year of the initial
commit in the rook/rook repository.
No more root user running.
Closes: https://github.com/rook/rook/issues/8734
Signed-off-by: Sébastien Han <seb@redhat.com>
We can pass bound_service_account_names with a comma separated list of
service accounts. Let's do this instead of remapping new values.
Earlier, we thought a single service account could be added per Vault
role and we were using other variables like
`VAULT_AUTH_KUBERNETES_ROOK_OPERATOR_ROLE` that we were remapping to
`VAULT_AUTH_KUBERNETES_ROLE` internal for the API calls to Vault.
Signed-off-by: Sébastien Han <seb@redhat.com>
The monitor list was not sorted, so each time we were reconciling, the
peer secret token will see its content updated with randomized
monitors. This would enter our predicate and trigger a reconcile.
Potentially an endless one, if the randomized list is already different.
Closes: https://github.com/rook/rook/issues/9076
Signed-off-by: Sébastien Han <seb@redhat.com>
Rook cluster-wide encryption can now use the native Kubernetes
authentication to interact with vault KMS instead of using the token
method.
Signed-off-by: Sébastien Han <seb@redhat.com>
When deploying a cluster-wde encrypted cluster we now use the rook
binary to execute some code to fetch the key encryption key.
Using cURL all the time to fetch the key has its limitations. The
incoming integration with Kubernetes Authentication through service
accounts is leading the usage of cURL to its end.
The logic is really complex and prone to errors. Re-implementing the
Go logic into Bash is not a viable option.
Also, using the lib in a binary allows us to keep a consistent behavior
throughout the life cycle of our code.
Signed-off-by: Sébastien Han <seb@redhat.com>
Previously we were merging the stderr even if it was empty, leading to
unmarshall errors.
The error simulation was done here
https://play.golang.org/p/Sk2yw9GUWNu.
Signed-off-by: Sébastien Han <seb@redhat.com>
Prior to ceph v16.2.7 the failover of the arbiter mon was
not supported. Now the new tiebreaker mon can be set during
the failover event and provide more dynamic stability to
the mon quorum if another node is available in the arbiter
zone.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When TLS is used and includes a caert, client key/cert, we need to copy
the content of the secret to a file in the operator's container
filesystem so that we can build the TLS config and thus the HTTP Client,
which reads those files.
Also, removing the files after each API call so they don't persist on
the filesystem forever.
Signed-off-by: Sébastien Han <seb@redhat.com>
Using a default value for CompressionMode to none effectively overrides
any values for Parameters. It is deprecated but still takes precedence.
Which means that in its previous form, Parameters was always ignored
since CompressionMode was always set to none when empty.
Signed-off-by: Sébastien Han <seb@redhat.com>
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.
If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.
Resolves#8778
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Prior, we were returning a nil map and thus the assignment for forced
deletion was not working since we were trying to assign on a nil map.
Signed-off-by: Sébastien Han <seb@redhat.com>
When proxying commands to the cmd-proxy container we don't need to build
the command line with the same flags as the operator. The cmd-proxy
container does not use any ceph config file and just relies on the
CEPH_ARGS environment variable in the container. So passing the same
args as the operator causes to fail since we don't have a ceph config
file in `/var/lib/rook/openshift-storage/openshift-storage.config` thus
the remote exec fails with:
```
global_init: unable to open config file from search list ...
```
Signed-off-by: Sébastien Han <seb@redhat.com>
We don't need to use tini.
We don't have anything in the rook operator that would
either create zombie processes (no threads) or use
exec (to fork). The Go binary has a really good
signal handling mechanism.
Closes: https://github.com/rook/rook/issues/8794
Signed-off-by: Sébastien Han <seb@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
The error message in UpdateNodeStatus regards the second argument
as node name. However, it is a PVC name in OSD on PVC.
Signed-off-by: Hiroya Onoe <onoehiroya@gmail.com>
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters
This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Passing struct by value essentially gives you a copy, so when modified
within a function, the scope is then reduced to that function. Using
pointers solves that you mutate the struct as many times as you want from
anywhere.
As a result, the auto-detection of the Vault KV backend was not working
correctly.
Also, added a ton of unit tests for Vault.
Signed-off-by: Sébastien Han <seb@redhat.com>
This handles the scenario where the OSDs have been created but not yet
started due to a wrong CR configuration.
For instance, when OSDs are encrypted and Vault is used to store
encryption keys, if the KV version is incorrect during the cluster
initialization the OSDs will fail to start and stay in CLBO until the
CR is updated again with the correct KV version so that it can start.
For this scenario, if the CRUSH map has no host registered yet it's fair
to assume the initialization broke and we need to fix it. So when don't
need to call ok-to-stop since it will always fail and eventually force
pass but let's not wait for nothing.
Signed-off-by: Sébastien Han <seb@redhat.com>
The mon config had two different implementations that have evolved
over the lifetime of the project. This is a simple refactor to remove
the SetConfig() option and stick with the MonStore as a single
implementation for updating the mon store.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Trying to disable mirroring on a cluster where mirroring is not enabled
results in an error during upgrades. Let's catch this error and ignore
it.
Funny enough the exec error resembles to:
```
Error ENOTSUP: Module 'mirroring' is not enabled (required by command 'fs snapshot mirror disable'): use `ceph mgr module enable mirroring` to enable it: exit status 95
```
So we get ENOTSUP which is 45 but exit status 95...
Closes: https://github.com/rook/rook/issues/8438
Signed-off-by: Sébastien Han <seb@redhat.com>