Commit Graph
399 Commits
Author SHA1 Message Date
Sébastien Han 7402c2cce6 osd: check if osd is ok-to-stop before removal
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.

So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-25 10:21:32 +01:00
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Sébastien Han c5783a77cf Merge pull request #9163 from y1r/add-context-k8sutil-node
core: add context parameter to k8sutil node
2021-11-15 16:25:13 +01:00
Satoru Takeuchi 07e1ca678a Merge pull request #9168 from cybozu-go/core-fix-unnecessary-option
core: remove unnecessary option
2021-11-15 22:50:08 +09:00
Yuichiro Ueno 4cc716a7ca core: add context parameter to k8sutil node
This commit adds context parameter to k8sutil node functions. By this,
we can handle cancellation during API call of node resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:45:59 +09:00
Satoru Takeuchi 77e7c249cc core: remove unnecessary option
rook command doesn't interpret `logtostderr` option. It's OK to just
remove this option because `capnslog` outputs all logs to stdout
by default. It's better to keep `AddGoFlagSet()` call because
some libraries might define their own flags with Go's `flag` package.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-11-15 11:53:45 +00:00
Yuichiro Ueno 0559977b8a core: add context parameter to k8sutil pod
This commit adds context parameter to k8sutil pod functions. By this, we
can handle cancellation during API call of pod resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 15:39:41 +09:00
Sébastien Han 16729e0e38 osd: use multiple service account for vault role
We can pass bound_service_account_names with a comma separated list of
service accounts. Let's do this instead of remapping new values.
Earlier, we thought a single service account could be added per Vault
role and we were using other variables like
`VAULT_AUTH_KUBERNETES_ROOK_OPERATOR_ROLE` that we were remapping to
`VAULT_AUTH_KUBERNETES_ROLE` internal for the API calls to Vault.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-04 15:15:46 +01:00
Sébastien Han 18a4047679 osd: add support for k8s with vault kms
Rook cluster-wide encryption can now use the native Kubernetes
authentication to interact with vault KMS instead of using the token
method.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-21 13:59:28 +02:00
Blaine Gardner 063c714c97 Merge pull request #9007 from BlaineEXE/remove-dynamic-clientset
ceph: get rid of dynamic clientset
2021-10-20 09:40:28 -06:00
Blaine Gardner 01e2feaef5 ceph: get rid of dynamic clientset
Stop using the dynamic clientset in favor of the controller-runtime
clientset.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-19 17:05:38 -06:00
Blaine Gardner d43327de2a docs: move purge osd to cluster namespace
The rook-ceph-purge-osd job should be run for a particular CephCluster
and should therefore be namespaced to the cluster.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-19 13:23:42 -06:00
Sébastien Han 4017f94464 osd: use rook binary to fetch key encryption key
When deploying a cluster-wde encrypted cluster we now use the rook
binary to execute some code to fetch the key encryption key.

Using cURL all the time to fetch the key has its limitations. The
incoming integration with Kubernetes Authentication through service
accounts is leading the usage of cURL to its end.
The logic is really complex and prone to errors. Re-implementing the
Go logic into Bash is not a viable option.

Also, using the lib in a binary allows us to keep a consistent behavior
throughout the life cycle of our code.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-19 17:29:25 +02:00
Travis Nielsen 87eee8afbc Merge pull request #8924 from leseb/rm-unused-client
core: remove unused clientset
2021-10-06 07:40:52 -06:00
Sébastien Han 723f144452 core: remove unused clientset
Removing ancient clientset which was not used anywhere any more.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-06 11:34:06 +02:00
Blaine Gardner 7586cea049 core: create TRACE_INSECURE log level
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.

If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.

Resolves #8778

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-01 12:09:19 -06:00
Sébastien Han 121c2987e3 ceph: stop using tini
We don't need to use tini.
We don't have anything in the rook operator that would
either create zombie processes (no threads) or use
exec (to fork). The Go binary has a really good
signal handling mechanism.

Closes: https://github.com/rook/rook/issues/8794
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-27 10:54:16 +02:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Hiroya Onoe 2ff5413b75 ceph: fix error message in UpdateNodeStatus
The error message in UpdateNodeStatus regards the second argument
as node name. However, it is a PVC name in OSD on PVC.

Signed-off-by: Hiroya Onoe <onoehiroya@gmail.com>
2021-09-02 02:43:16 +00:00
Blaine Gardner a1814af1d9 ceph: remove NFS and Cassandra operator code
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-08-31 14:07:02 -06:00
Sébastien Han 41e915d411 Merge pull request #8493 from leseb/admission-controller
ceph: move the admission webhook to the operator
2021-08-25 09:30:54 +02:00
Sébastien Han 656dd0f334 ceph: move the admission webhook to the operator
Our admission webhooks will now run as part of the Operator container
and not an additional deployment. This has the advantage of consuming
fewer resources in the cluster and not having to manage affinities and
tolerations. This only drawback is that the Secret containing the
certificates is not mounted anymore and the content needs to be written
inside the Operator. This is not practical since we also need to watch
for the Secret content to change. Meaning that the certificates have
been renewed and the webhook server needs to use them.
A new approach is on its way to hopefully simplify this last issue and
implement a watcher for the Secret.
In the meantime, users need to use the cert-manager or renew
certificates manually. Additionally, they must update the
ValidatingWebhookConfiguration object with the new CA bundle.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-24 19:07:04 +02:00
Sébastien Han 579fc27783 ceph: embed ceph-csi templates in rook binary
Thanks to Golang 1.16, we can now embed files in the Go binary. This
means we don't need to add the CSI templates files to the container
image. They are added in the Go binary at build time.

The existing location of the template must be in the package calling it.
So they moved to pkg/operator/ceph/csi/template. All the files have been
symlinked back to cluster/examples/kubernetes/ceph/csi/template.

Closes: https://github.com/rook/rook/issues/7609
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-24 16:22:20 +02:00
Sébastien Han 07d14a93a2 ceph: remove cli unused flags
I don't know why these flags are there but it's not like we run the
operator with them and with a different value.
So removing for clarity.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-06 11:05:20 +02:00
Sébastien Han 8fdb9fe2b9 ceph: remove old network types
Those types are legacy, not really used anywhere and not needed anymore.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-04 10:56:01 +02:00
Satoru Takeuchi 89b6a6028c ceph: add an option to preserve pvc in osd purge job
Sometimes we want to investigate a PVC for removed OSD.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-29 14:20:06 +00:00
Travis Nielsen 107f89260d ceph: reconcile mgr services with every reconcile
When there are multiple mgr daemons, the mgr sidecar owns reconciling
the services for the active mgr. However, for efficiency the sidecar
only reconciles the services if the active mgr changes. Now the operator
will also reconcile the mgr services to ensure that services are re-created
after being deleted, or otherwise in an incorrect state.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-20 07:23:58 -06:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
subhamkraiandSébastien Han 4241ce7e66 ceph: operator pod killed due to resource limit
when operator pod resources are set(which are
default in case of helm), operator pod are
killed due to using concurrency while creating
object store.

this commit checks if operator resources are set
or not. if set then we'll *not* create object store
in concurrency or if not set then we'll create in
concurrency.

Closes: https://github.com/rook/rook/issues/8149
Co-authored-by: Sébastien Han <seb@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2021-07-02 20:18:58 +05:30
Andy Bursavich 11f0add706 build: update golangci-lint version and nolint comment format
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:50:04 -07:00
Travis Nielsen ca9180d22e Merge pull request #8098 from degorenko/parse-devices
ceph: do not fail prepareOSD job if devices are not passed
2021-06-10 11:41:08 -06:00
Blaine Gardner 21e290e003 ceph: implement dependencies for CephCluster
Implement the first step of `design/ceph/resource-dependencies.md` to
add dependency checking when deleting a CephCluster.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-10 10:06:10 -06:00
Denis Egorenko fa2c276cc1 ceph: do not fail prepareOSD job if devices are not passed
If devices are not specified for node return empty desired devices
list and do not fail prepareOSD job.

Closes: #8097
Signed-off-by: Denis Egorenko <degorenko@mirantis.com>
2021-06-10 15:55:10 +04:00
Travis Nielsen eba91a4bd6 ceph: remove obsolete references to filestore
Filestore support has been long gone with only bluestore
currently supported.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-04 15:50:20 -06:00
Shachar Sharon 9766e9b8fb ceph: allow passing 'osd-crush-initial-weight'
Ceph support the option '--osd-crush-initial-weight' upon OSD start,
which sets an explicit weight (in TiB units) to specific OSD. Allow
passing this option all the way from the user (similar to
'DeviceClass'), for the special case where end users wants it cluster
to have non-even balance over specific OSDs (e.g., one of the OSDs is
placed over a partition alongside OS-partition).

ROOK issue: https://github.com/rook/rook/issues/7448

Signed-off-by: Shachar Sharon <ssharon@redhat.com>
2021-04-20 20:50:54 +03:00
Travis Nielsen 64e28af741 ceph: allow flex driver and discovery to be enabled with configmap
For testing purposes, we need to configure the flex driver
and discovery daemon with the operator settings configmap.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 08:39:26 -06:00
Travis Nielsen 933113eac8 ceph: remove obsolete storeType setting
All clusters are configured with bluestore as the store type
therefore we can remove the settings where mentioned in the
docs and dead code.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 08:39:25 -06:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Travis Nielsen 38efe8ceba ceph: get active mgr with new pacific command
In Pacific the simple command to retrieve the active mgr is with
ceph mgr stat instead of retrieving the full ceph mgr dump.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:37:18 -07:00
Travis Nielsen 388ff3e78b ceph: start sidecar to monitor active mgr
The mgr daemon may be failed over by ceph if the active mgr is not
responding and the standby mgr is available. If the active mgr changes
the services for the dashboard and metrics will be updated with a
label selector for the new active mgr. The services cannot direct
traffic to the standby mgr or else they will be incorrectly redirected
to the active mgr.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:37:17 -07:00
Blaine Gardner efa7d66d45 ceph: remove driveGroup support
Remove features and design that supports adding OSDs to Ceph clusters
via `spec:driveGroups`. Update the Ceph upgrade doc that informs users
who currently use Drive Groups (we believe there are none of these
users) how to migrate to using the `spec:storage` config.

Resolves https://github.com/rook/rook/issues/7275

Revert "ceph: fix drive group deployment failure"
This reverts commit 76f1d9944e.

Revert "ceph: osd: add drive groups spec to cluster CR"
This reverts commit 7117fc12b7.

Revert "design: ceph orchestrator module add/remove OSDs"
This reverts commit 178187d035.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-02-25 17:08:47 -07:00
Travis Nielsen 291bf9e061 ceph: enforce portable osds in same topology as osd prepare job
The osd prepare job detects the failure domain for the OSD by querying
the node topology labels. This topology is then assigned to the OSD
daemon in the CRUSH map. Previously, the affinity was required to be
set in the cluster CR, but it was very difficult to get right. Now the
operator will enforce the correct topology label on the OSD daemon
nodeAffinity by setting the label of the lowest topology in the hierarcy.

For example, if there are region, zone, and rack labels, the rack label
would be used to set the node affinity for the OSD daemon. If an
OSD prepare job is run in rack1, the corresponding OSD daemon will
have node affinity to rack1 to ensure the same topology. Previously,
the OSD could have ended up in rack2 unless the storageClassDeviceSet
placement was very carefully crafted.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-02-18 12:49:37 -07:00
Travis Nielsen 043f947321 build: remove yugabytedb operator
The YugabyteDB team has decided to move support of their operator
to a new repo at https://github.com/yugabyte/yugabyte-operator.
The operator in Rook is no longer needed. Further usage of
the YugabyteDB operator is recommended at the new location.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-02-06 00:03:41 -07:00
Travis Nielsen 4b43e7afa5 ceph: allow removal of arbitrary osds on pvcs and simplify pvc names
OSDs on PVCs previously were always replaced with PVCs of a given name.
For on-prem scenarios where the OSDs might need to be removed instead
of replaced, the operator now allows holes to exist in the PVC index
names as long as there are a sufficient number of PVCs to meet the
criteria for the deviceSet.count.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-26 14:38:12 -07:00
Travis Nielsen 53ed11f15b build: remove the edgefs operator from rook
The EdgeFS operator has been deprecated for some time in Rook.
If the replacement is added back to Rook it can be completed
according to the new guidelines in the documentation.
https://rook.io/docs/rook/master/storage-providers.html

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-25 17:51:24 -07:00
Travis Nielsen 2294851ca3 build: remove the cockroachdb operator from rook
The cockroachDB operator has not had community support in Rook.
Therefore, the time has come to deprecate and remove it.
If the sources are still needed, there is always git history.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-20 17:28:58 -07:00
Travis Nielsenandshenjiatong a0a27bc426 ceph: apply deviceClass properly to new osds
The deviceClass property was being ignored when creating the
non-pvc OSDs. Now the deviceClass will be specified as a property
for individual devices, all devices on a node, or all OSDs in the
cluster, depending on the level where the config is applied in the
cluster CR.

Co-authored-by: shenjiatong <yshxxsjt715@gmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-11 10:16:00 -07:00
Sébastien Han ad24990473 ceph: ability to abort orchestration
We can now prioritize orchestrations on certain event. Today only two
events will cancel on-going orchestrations (if any):

* request for cluster deletion
* request for cluster upgrade

If one of the two are caught by the watcher we will cancel the on-going
orchestration.
For that we implemented a simple approach based on check points, where
we will check for a cancellation request in certain part of the
orchestration. Mainly before each mons/mgr/osds orchestration loops.

This solution is not perfect, but we are waiting for the
controller-runtime to release its 0.7 version which will embed context
support. With that we will be able to cancel reconciles more precisely
and rapidly.

Operator log example:

```
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-25 12:59:12.986264 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.776947 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
                Image:            "ceph/ceph:v15.2.5",
-               AllowUnsupported: true,
+               AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:33.777039 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.785088 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:33.788626 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.5...
2020-11-25 13:07:35.280789 I | ceph-cluster-controller: detected ceph image version: "15.2.5-0 octopus"
2020-11-25 13:07:35.280806 I | ceph-cluster-controller: validating ceph version from provided image
2020-11-25 13:07:35.285888 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.287828 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.288082 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:35.621625 I | ceph-cluster-controller: cluster "rook-ceph": version "15.2.5-0 octopus" detected for image "ceph/ceph:v15.2.5"
2020-11-25 13:07:35.642688 I | op-mon: start running mons
2020-11-25 13:07:35.646323 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.654070 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:35.868253 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.868573 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:37.074353 I | op-mon: targeting the mon count 3
2020-11-25 13:07:38.153435 I | op-mon: checking for basic quorum with existing mons
2020-11-25 13:07:38.178029 I | op-mon: mon "a" endpoint is [v2:10.107.242.49:3300,v1:10.107.242.49:6789]
2020-11-25 13:07:38.670191 I | op-mon: mon "b" endpoint is [v2:10.109.71.30:3300,v1:10.109.71.30:6789]
2020-11-25 13:07:39.477820 I | op-mon: mon "c" endpoint is [v2:10.98.93.224:3300,v1:10.98.93.224:6789]
2020-11-25 13:07:39.874094 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:40.467999 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:40.469733 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.071710 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:41.078903 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.125233 I | op-mon: deployment for mon rook-ceph-mon-a already exists. updating if needed
2020-11-25 13:07:41.327778 I | op-k8sutil: updating deployment "rook-ceph-mon-a" after verifying it is safe to stop
2020-11-25 13:07:41.327895 I | op-mon: checking if we can stop the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045644 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-a"
2020-11-25 13:07:44.045706 I | op-mon: checking if we can continue the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045740 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:44.109159 I | op-mon: mons running: [a b c]
2020-11-25 13:07:44.474596 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:44.478565 I | op-mon: deployment for mon rook-ceph-mon-b already exists. updating if needed
2020-11-25 13:07:44.493374 I | op-k8sutil: updating deployment "rook-ceph-mon-b" after verifying it is safe to stop
2020-11-25 13:07:44.493403 I | op-mon: checking if we can stop the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135524 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-b"
2020-11-25 13:07:47.135542 I | op-mon: checking if we can continue the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135551 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:47.148820 I | op-mon: mons running: [a b c]
2020-11-25 13:07:47.445946 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:47.448991 I | op-mon: deployment for mon rook-ceph-mon-c already exists. updating if needed
2020-11-25 13:07:47.462041 I | op-k8sutil: updating deployment "rook-ceph-mon-c" after verifying it is safe to stop
2020-11-25 13:07:47.462060 I | op-mon: checking if we can stop the deployment rook-ceph-mon-c
2020-11-25 13:07:48.853118 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff=  v1.ClusterSpec{
        CephVersion: v1.CephVersionSpec{
-               Image:            "ceph/ceph:v15.2.5",
+               Image:            "ceph/ceph:v15.2.6",
                AllowUnsupported: false,
        },
        DriveGroups: nil,
        Storage:     {UseAllNodes: true, Selection: {UseAllDevices: &true}},
        ... // 20 identical fields
  }
2020-11-25 13:07:48.853140 I | ceph-cluster-controller: upgrade requested, cancelling any ongoing orchestration
2020-11-25 13:07:50.119584 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-c"
2020-11-25 13:07:50.119606 I | op-mon: checking if we can continue the deployment rook-ceph-mon-c
2020-11-25 13:07:50.119619 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.130860 I | op-mon: mons running: [a b c]
2020-11-25 13:07:50.431341 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:50.431361 I | op-mon: mons created: 3
2020-11-25 13:07:50.734156 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.745763 I | op-mon: mons running: [a b c]
2020-11-25 13:07:51.045108 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:51.054497 E | ceph-cluster-controller: failed to reconcile. failed to reconcile cluster "rook-ceph": failed to configure local ceph cluster: failed to create cluster: CANCELLING CURRENT ORCHESTATION
2020-11-25 13:07:52.055208 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:52.070690 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:52.088979 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.6...
2020-11-25 13:07:53.904811 I | ceph-cluster-controller: detected ceph image version: "15.2.6-0 octopus"
2020-11-25 13:07:53.904862 I | ceph-cluster-controller: validating ceph version from provided image
```

Closes: https://github.com/rook/rook/issues/6587
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-25 15:32:04 +01:00
Arun Kumar Mohan ded16f779d ceph: changes for 'sigs.k8s.io/sig-storage-lib-external-provisioner/v6'
Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:03 +05:30