This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
The error message in UpdateNodeStatus regards the second argument
as node name. However, it is a PVC name in OSD on PVC.
Signed-off-by: Hiroya Onoe <onoehiroya@gmail.com>
Our admission webhooks will now run as part of the Operator container
and not an additional deployment. This has the advantage of consuming
fewer resources in the cluster and not having to manage affinities and
tolerations. This only drawback is that the Secret containing the
certificates is not mounted anymore and the content needs to be written
inside the Operator. This is not practical since we also need to watch
for the Secret content to change. Meaning that the certificates have
been renewed and the webhook server needs to use them.
A new approach is on its way to hopefully simplify this last issue and
implement a watcher for the Secret.
In the meantime, users need to use the cert-manager or renew
certificates manually. Additionally, they must update the
ValidatingWebhookConfiguration object with the new CA bundle.
Signed-off-by: Sébastien Han <seb@redhat.com>
Thanks to Golang 1.16, we can now embed files in the Go binary. This
means we don't need to add the CSI templates files to the container
image. They are added in the Go binary at build time.
The existing location of the template must be in the package calling it.
So they moved to pkg/operator/ceph/csi/template. All the files have been
symlinked back to cluster/examples/kubernetes/ceph/csi/template.
Closes: https://github.com/rook/rook/issues/7609
Signed-off-by: Sébastien Han <seb@redhat.com>
I don't know why these flags are there but it's not like we run the
operator with them and with a different value.
So removing for clarity.
Signed-off-by: Sébastien Han <seb@redhat.com>
When there are multiple mgr daemons, the mgr sidecar owns reconciling
the services for the active mgr. However, for efficiency the sidecar
only reconciles the services if the active mgr changes. Now the operator
will also reconcile the mgr services to ensure that services are re-created
after being deleted, or otherwise in an incorrect state.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.
So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.
Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.
It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.
Signed-off-by: Sébastien Han <seb@redhat.com>
when operator pod resources are set(which are
default in case of helm), operator pod are
killed due to using concurrency while creating
object store.
this commit checks if operator resources are set
or not. if set then we'll *not* create object store
in concurrency or if not set then we'll create in
concurrency.
Closes: https://github.com/rook/rook/issues/8149
Co-authored-by: Sébastien Han <seb@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
Implement the first step of `design/ceph/resource-dependencies.md` to
add dependency checking when deleting a CephCluster.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
If devices are not specified for node return empty desired devices
list and do not fail prepareOSD job.
Closes: #8097
Signed-off-by: Denis Egorenko <degorenko@mirantis.com>
Ceph support the option '--osd-crush-initial-weight' upon OSD start,
which sets an explicit weight (in TiB units) to specific OSD. Allow
passing this option all the way from the user (similar to
'DeviceClass'), for the special case where end users wants it cluster
to have non-even balance over specific OSDs (e.g., one of the OSDs is
placed over a partition alongside OS-partition).
ROOK issue: https://github.com/rook/rook/issues/7448
Signed-off-by: Shachar Sharon <ssharon@redhat.com>
For testing purposes, we need to configure the flex driver
and discovery daemon with the operator settings configmap.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
All clusters are configured with bluestore as the store type
therefore we can remove the settings where mentioned in the
docs and dead code.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
In Pacific the simple command to retrieve the active mgr is with
ceph mgr stat instead of retrieving the full ceph mgr dump.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The mgr daemon may be failed over by ceph if the active mgr is not
responding and the standby mgr is available. If the active mgr changes
the services for the dashboard and metrics will be updated with a
label selector for the new active mgr. The services cannot direct
traffic to the standby mgr or else they will be incorrectly redirected
to the active mgr.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Remove features and design that supports adding OSDs to Ceph clusters
via `spec:driveGroups`. Update the Ceph upgrade doc that informs users
who currently use Drive Groups (we believe there are none of these
users) how to migrate to using the `spec:storage` config.
Resolves https://github.com/rook/rook/issues/7275
Revert "ceph: fix drive group deployment failure"
This reverts commit 76f1d9944e.
Revert "ceph: osd: add drive groups spec to cluster CR"
This reverts commit 7117fc12b7.
Revert "design: ceph orchestrator module add/remove OSDs"
This reverts commit 178187d035.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The osd prepare job detects the failure domain for the OSD by querying
the node topology labels. This topology is then assigned to the OSD
daemon in the CRUSH map. Previously, the affinity was required to be
set in the cluster CR, but it was very difficult to get right. Now the
operator will enforce the correct topology label on the OSD daemon
nodeAffinity by setting the label of the lowest topology in the hierarcy.
For example, if there are region, zone, and rack labels, the rack label
would be used to set the node affinity for the OSD daemon. If an
OSD prepare job is run in rack1, the corresponding OSD daemon will
have node affinity to rack1 to ensure the same topology. Previously,
the OSD could have ended up in rack2 unless the storageClassDeviceSet
placement was very carefully crafted.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The YugabyteDB team has decided to move support of their operator
to a new repo at https://github.com/yugabyte/yugabyte-operator.
The operator in Rook is no longer needed. Further usage of
the YugabyteDB operator is recommended at the new location.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
OSDs on PVCs previously were always replaced with PVCs of a given name.
For on-prem scenarios where the OSDs might need to be removed instead
of replaced, the operator now allows holes to exist in the PVC index
names as long as there are a sufficient number of PVCs to meet the
criteria for the deviceSet.count.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The cockroachDB operator has not had community support in Rook.
Therefore, the time has come to deprecate and remove it.
If the sources are still needed, there is always git history.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The deviceClass property was being ignored when creating the
non-pvc OSDs. Now the deviceClass will be specified as a property
for individual devices, all devices on a node, or all OSDs in the
cluster, depending on the level where the config is applied in the
cluster CR.
Co-authored-by: shenjiatong <yshxxsjt715@gmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We can now prioritize orchestrations on certain event. Today only two
events will cancel on-going orchestrations (if any):
* request for cluster deletion
* request for cluster upgrade
If one of the two are caught by the watcher we will cancel the on-going
orchestration.
For that we implemented a simple approach based on check points, where
we will check for a cancellation request in certain part of the
orchestration. Mainly before each mons/mgr/osds orchestration loops.
This solution is not perfect, but we are waiting for the
controller-runtime to release its 0.7 version which will embed context
support. With that we will be able to cancel reconciles more precisely
and rapidly.
Operator log example:
```
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-24 13:54:59.499719 I | op-mon: parsing mon endpoints: a=10.109.126.120:6789
2020-11-25 12:59:12.986264 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.776947 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff= v1.ClusterSpec{
CephVersion: v1.CephVersionSpec{
Image: "ceph/ceph:v15.2.5",
- AllowUnsupported: true,
+ AllowUnsupported: false,
},
DriveGroups: nil,
Storage: {UseAllNodes: true, Selection: {UseAllDevices: &true}},
... // 20 identical fields
}
2020-11-25 13:07:33.777039 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:33.785088 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:33.788626 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.5...
2020-11-25 13:07:35.280789 I | ceph-cluster-controller: detected ceph image version: "15.2.5-0 octopus"
2020-11-25 13:07:35.280806 I | ceph-cluster-controller: validating ceph version from provided image
2020-11-25 13:07:35.285888 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.287828 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.288082 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:35.621625 I | ceph-cluster-controller: cluster "rook-ceph": version "15.2.5-0 octopus" detected for image "ceph/ceph:v15.2.5"
2020-11-25 13:07:35.642688 I | op-mon: start running mons
2020-11-25 13:07:35.646323 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:35.654070 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:35.868253 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:35.868573 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:37.074353 I | op-mon: targeting the mon count 3
2020-11-25 13:07:38.153435 I | op-mon: checking for basic quorum with existing mons
2020-11-25 13:07:38.178029 I | op-mon: mon "a" endpoint is [v2:10.107.242.49:3300,v1:10.107.242.49:6789]
2020-11-25 13:07:38.670191 I | op-mon: mon "b" endpoint is [v2:10.109.71.30:3300,v1:10.109.71.30:6789]
2020-11-25 13:07:39.477820 I | op-mon: mon "c" endpoint is [v2:10.98.93.224:3300,v1:10.98.93.224:6789]
2020-11-25 13:07:39.874094 I | op-mon: saved mon endpoints to config map map[csi-cluster-config-json:[{"clusterID":"rook-ceph","monitors":["10.107.242.49:6789","10.109.71.30:6789","10.98.93.224:6789"]}] data:a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789 mapping:{"node":{"a":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"b":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"},"c":{"Name":"minikube","Hostname":"minikube","Address":"192.168.39.3"}}} maxMonId:2]
2020-11-25 13:07:40.467999 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:40.469733 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.071710 I | cephclient: writing config file /var/lib/rook/rook-ceph/rook-ceph.config
2020-11-25 13:07:41.078903 I | cephclient: generated admin config in /var/lib/rook/rook-ceph
2020-11-25 13:07:41.125233 I | op-mon: deployment for mon rook-ceph-mon-a already exists. updating if needed
2020-11-25 13:07:41.327778 I | op-k8sutil: updating deployment "rook-ceph-mon-a" after verifying it is safe to stop
2020-11-25 13:07:41.327895 I | op-mon: checking if we can stop the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045644 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-a"
2020-11-25 13:07:44.045706 I | op-mon: checking if we can continue the deployment rook-ceph-mon-a
2020-11-25 13:07:44.045740 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:44.109159 I | op-mon: mons running: [a b c]
2020-11-25 13:07:44.474596 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:44.478565 I | op-mon: deployment for mon rook-ceph-mon-b already exists. updating if needed
2020-11-25 13:07:44.493374 I | op-k8sutil: updating deployment "rook-ceph-mon-b" after verifying it is safe to stop
2020-11-25 13:07:44.493403 I | op-mon: checking if we can stop the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135524 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-b"
2020-11-25 13:07:47.135542 I | op-mon: checking if we can continue the deployment rook-ceph-mon-b
2020-11-25 13:07:47.135551 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:47.148820 I | op-mon: mons running: [a b c]
2020-11-25 13:07:47.445946 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:47.448991 I | op-mon: deployment for mon rook-ceph-mon-c already exists. updating if needed
2020-11-25 13:07:47.462041 I | op-k8sutil: updating deployment "rook-ceph-mon-c" after verifying it is safe to stop
2020-11-25 13:07:47.462060 I | op-mon: checking if we can stop the deployment rook-ceph-mon-c
2020-11-25 13:07:48.853118 I | ceph-cluster-controller: CR has changed for "rook-ceph". diff= v1.ClusterSpec{
CephVersion: v1.CephVersionSpec{
- Image: "ceph/ceph:v15.2.5",
+ Image: "ceph/ceph:v15.2.6",
AllowUnsupported: false,
},
DriveGroups: nil,
Storage: {UseAllNodes: true, Selection: {UseAllDevices: &true}},
... // 20 identical fields
}
2020-11-25 13:07:48.853140 I | ceph-cluster-controller: upgrade requested, cancelling any ongoing orchestration
2020-11-25 13:07:50.119584 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mon-c"
2020-11-25 13:07:50.119606 I | op-mon: checking if we can continue the deployment rook-ceph-mon-c
2020-11-25 13:07:50.119619 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.130860 I | op-mon: mons running: [a b c]
2020-11-25 13:07:50.431341 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:50.431361 I | op-mon: mons created: 3
2020-11-25 13:07:50.734156 I | op-mon: waiting for mon quorum with [a b c]
2020-11-25 13:07:50.745763 I | op-mon: mons running: [a b c]
2020-11-25 13:07:51.045108 I | op-mon: Monitors in quorum: [a b c]
2020-11-25 13:07:51.054497 E | ceph-cluster-controller: failed to reconcile. failed to reconcile cluster "rook-ceph": failed to configure local ceph cluster: failed to create cluster: CANCELLING CURRENT ORCHESTATION
2020-11-25 13:07:52.055208 I | ceph-cluster-controller: reconciling ceph cluster in namespace "rook-ceph"
2020-11-25 13:07:52.070690 I | op-mon: parsing mon endpoints: a=10.107.242.49:6789,b=10.109.71.30:6789,c=10.98.93.224:6789
2020-11-25 13:07:52.088979 I | ceph-cluster-controller: detecting the ceph image version for image ceph/ceph:v15.2.6...
2020-11-25 13:07:53.904811 I | ceph-cluster-controller: detected ceph image version: "15.2.6-0 octopus"
2020-11-25 13:07:53.904862 I | ceph-cluster-controller: validating ceph version from provided image
```
Closes: https://github.com/rook/rook/issues/6587
Signed-off-by: Sébastien Han <seb@redhat.com>
The discovery daemon is not needed in most scenarios, therefore we disable it
by default. More and more clusters are moving to the cluster-on-pvc scenario
which certainly does not need the local discovery. Even where clusters are not
running on PVCs, the discovery is not needed since the device discovery is again
performed in the osd prepare job.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.
By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.
To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.
The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.
Fixes#4993.
Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>
this commit handle golangci-lint linter errcheck.
`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases
To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`
Signed-off-by: subhamkrai <srai@redhat.com>
this commit will enable one more linter ineffassign
in golangci-lint.
This linter throws an error when variable is assigned and never used.
`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
Drive Groups will fail to deploy due to a message: `failed to parse
device list (""): failed to JSON unmarshal configured devices (""):
unexpected end of JSON input`
This is due to trying to JSON unmarshal an empty string which does not
contain any devices to provision when Drive Groups are specified. Fix so
that an empty string represents no non-Drive Group devices.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
If an OSD is down and needs to be removed from the cluster,
a job can be run that will remove the OSD deployment
and purge the OSD from ceph. If the OSD is still up,
the purge will be rejected.
To purge multiple OSDs at the same time, the OSD IDs can be
specified as a comma-separated list.
Signed-off-by: Servesha Dudhgaonkar <sdudhgao@redhat.com>
as the kubernetes 1.13 is EOL and there is no major
functionalities available in 1.13 (resize,snapshot,clone
metrics etc) we are removing the support for kubernetes
for the same.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
We can now encrypted OSD device that were provisioned via a storage
class using the PV interface.
The encryption works at the storageClassDeviceSets level, which means we
can have encrypted and non-encrypted sets.
Using the new key `encrypted` we can turn it on.
Signed-off-by: Sébastien Han <seb@redhat.com>
using defer for closing file which are open
for writing is not safe. so closing file again
following below steps:
1. open files
2. defer file.close()
3. write
4.file.close()
these will make sure files are closed.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
this commit suppress the gosec errors for
g204: Audit use of command execution.
g304: File path provided as taint input.
g101: Look for hard coded credentials.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
To ensure a file handle is closed, we defer the close command
so it is guaranteed to run when the method returns. The closing
of the file handle is not going to fail in our usage since we aren't
using the SetDeadline on the files that would cancel a request
and return an error.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Now, we can not only cleanup monitor data, logs and crashes but the
disks too. As part of the cleanupPolicy CR spec, we have a new setting
called sanitizeDisk which holds more details:
* method: indicates if the entire disk should be sanitized or simply ceph's metadata.
Possible choices are 'complete' or 'quick' (default)
* dataSource: indicate where to get random bytes from to write on the disk.
Possible choices are 'zero' (default) or 'random'
Using random sources will consume entropy from the system and will take much more time then the zero source
* iteration: overwrite N times instead of the default (1). Takes an integer value
See the documentation for more details.
Signed-off-by: Sébastien Han <seb@redhat.com>