Commit Graph
182 Commits
Author SHA1 Message Date
Humble Chirammal e5f5d9be9f csi: mount host's /etc/selinux in node plugins
This commit introduces a new configuration option for
ceph csi driver to enable hostpath mounting of /etc/selinux
directory from the cluster node where csi plugin pods are
running, which inturn help the csi driver to specify
selinux-related mount options like context.

Ref# https://github.com/ceph/ceph-csi/issues/2295

The default value for this configuration is true and if cluster
nodes are running without selinux enabled, an admin can deploy
csi pods by specifying this option to `false` which skip the
host path mounting for the csi pods.

Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
2021-12-01 11:13:26 +05:30
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Sébastien Han d5643b439d Merge pull request #9161 from y1r/add-context-k8sutil-daemonset
core: add context parameter to k8sutil daemonset
2021-11-15 15:14:44 +01:00
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
Yuichiro Ueno e278812cc1 core: add context parameter to k8sutil daemonset
This commit adds context parameter to k8sutil daemonset functions. By
this, we can handle cancellation during API call of daemonset resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 15:09:31 +09:00
Yuichiro Ueno 0b575703c7 core: add context parameter to k8sutil deployment
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 14:58:13 +09:00
Rakshith R c1ef189b91 ceph: apply csi provisioner node-affinity to csi version check job
This commit makes sure csi provisioner node-affinity is applied
to the csi version check job as well, similar to the how
the existing csi provisioner toleration is applied to the job.

Fixes: #8323

Signed-off-by: Rakshith R <rar@redhat.com>
2021-10-13 15:18:23 +05:30
Renan Campos c4e70933f2 ceph: csi controller to use constants to fetch operator configmap
The CSI controller loop watches ConfigMaps and CephClusters. When a
CephCluster triggers the reconcile loop, the Get call for the ConfigMap
will fail since the name in the request object is that of the
CephCluster. Setting the name and namespace to the constants for the
ConfigMap name and the operator namespace resolves this.

Closes: https://github.com/rook/rook/issues/8958
Signed-off-by: Renan Campos <rcampos@redhat.com>
2021-10-12 09:30:49 -04:00
Humble Chirammal 966a88ea81 ceph: update the csi sidecar versions in templates and documentation
Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
2021-09-24 18:25:12 +05:30
Travis Nielsen f5c1543ddf build: update min k8s version to 1.16
In Rook v1.8 the min version of K8s supported is updated to 1.16.
Users running on older versions of K8s are recommended to update
to 1.16 or newer before updating to Rook v1.8.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:56 -06:00
Travis Nielsen 0ec1303606 Merge pull request #8729 from Madhu-1/fix-8153
ceph: make provisioner replicas configurable
2021-09-22 11:00:00 -06:00
Madhu Rajanna 95775fd445 ceph: modify CephFS provisioner permission
As like RBD, CephFS provisioner pod need not to
run as privileged. as its not doing any operation
like plugin pods which does mounting and unmounting
removing the permissions for the same.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-09-22 13:30:05 +05:30
Madhu Rajanna ed5f281a74 ceph: make provisioner replicas configurable
added new option to set the provisioner replicas.
with this new option the user/admin can choose
how many replicas he want for provisioner pod if
number of nodes is greater than 1.

fixes #8153

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-09-22 11:24:53 +05:30
Humble Chirammal 731b0f5274 ceph: lift minimum supported version of ceph csi to v3.3.0
With the release of Ceph CSI 3.4.0, Ceph CSI project came up
with a new support policy where we support only versions >= 3.3.0

The supported window of Ceph CSI versions is known as "N.(x-1)":
 (N (Latest major release) . (x (Latest minor release) - 1)).

For example, if Ceph CSI latest major version is 3.4.0 today,
support is provided for the versions above 3.3.0.
If users are running an unsupported Ceph CSI version, they will be
asked to upgrade when requesting support for the cluster.

This PR lift the minimum supported version of Ceph CSI to 3.3.0
in this repo.

Fix https://github.com/rook/rook/issues/8709

Ref #
https://github.com/ceph/ceph-csi/releases/tag/v3.4.0
https://github.com/ceph/ceph-csi/#known-to-work-co-platforms

Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
2021-09-22 10:16:43 +05:30
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 68a4bc2d1b ceph: remove unnecessary package
We don't need to use github.com/ghodss/yaml since
"k8s.io/apimachinery/pkg/util/yaml" provides the same functionality and
we already import it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:10:47 +02:00
Santosh Pillai 3f8abec403 ceph: add ClusterID and PoolID mappings between local and peer cluster
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters

This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-08-31 20:08:56 +05:30
Travis Nielsen 9a96afab05 Merge pull request #8582 from Madhu-1/fix-csidriver-panic
ceph: fix panic when recreating the csidriver object
2021-08-24 09:05:31 -06:00
Sébastien Han 579fc27783 ceph: embed ceph-csi templates in rook binary
Thanks to Golang 1.16, we can now embed files in the Go binary. This
means we don't need to add the CSI templates files to the container
image. They are added in the Go binary at build time.

The existing location of the template must be in the package calling it.
So they moved to pkg/operator/ceph/csi/template. All the files have been
symlinked back to cluster/examples/kubernetes/ceph/csi/template.

Closes: https://github.com/rook/rook/issues/7609
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-24 16:22:20 +02:00
Madhu Rajanna d246dde443 ceph: remove variable declaration
remove csiDriver and csiClient variables
creation and reuse them from the method
receivers.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-08-24 19:43:18 +05:30
Madhu Rajanna e2dc418371 ceph: fix logging of csidriver
log the successful message about
starting CSIDriver after creating
the daemonset and deployment objects.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-08-24 19:41:04 +05:30
Madhu Rajanna d48f67a82c ceph: fix panic in reCreateCSIDriverInfo
Initialize the client and the csidriver object
before calling the reCreateCSIDriverInfo
function.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-08-24 12:15:48 +05:30
Humble Chirammal 8034a4ef98 ceph: set default FsGroupChangePolicy value to 'None'
The `None` value Indicates that volumes will be mounted
with no modifications, as the CSI volume driver does not
support these operations. While volumes are provisioned
by the CephFS CSI driver the global permissions are set
on the volume and we dont expect the Fsgroup policy or
check from CO side  to play a role here.

Ref #ceph/ceph-csi/../internal/cephfs/nodeserver.go#L190

```
!csicommon.MountOptionContains(fuseMountOptions, readOnly) {
		// #nosec - allow anyone to write inside the stagingtarget path
		err = os.Chmod(stagingTargetPath, 0o777)
```

The current default value ie `ReadWriteOnceWithFSType` cause
volumes to be examined to determine if volume ownership and permissions
should be modified to match the pod's security policy. Changes could occur
if the fsType is defined and the persistent volume's accessModes contains
ReadWriteOnce.

Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
2021-08-03 16:54:10 +05:30
Madhu Rajanna b556dbf109 ceph: update cephcsi to v3.4.0
This PR updates the required RBAC, templates,
CSI image version and examples for new
cephcsi v3.4.0 release.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-08-02 10:22:54 +05:30
Yug d513dc4bc5 ceph: update CSIDriver object from betav1 to v1
Update the CSIDriver object to v1 from betav1 as it will be
deprecated in v1.19+ and become unavailable in v1.22+

Signed-off-by: Yug <yuggupta27@gmail.com>
2021-07-02 10:19:53 +05:30
Rakshith R 2baf50771e ceph: update csi node-driver-registrar to latest release
This commit updates csi node-driver-registrar to latest
release v2.2.0.
https://github.com/kubernetes-csi/node-driver-registrar/releases

Signed-off-by: Rakshith R <rar@redhat.com>
2021-06-24 10:22:51 +05:30
Rakshith R 66c9ccab86 ceph: read and validate CSI params in Go routine
This commit moves SetParams and ValidateCSIParam func
into go routine/lock, since it modifies/reads global CSIParam
variable.

Signed-off-by: Rakshith R <rar@redhat.com>
2021-06-21 15:44:54 +05:30
Madhu Rajanna 08997434ac ceph: update csi sidecar images to latest release
This commit updates all the kubernetes sidecar
images to latest available release version.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-06-15 08:09:27 +05:30
Rakshith R f489aba8f2 ceph: set CSI_ENABLE_HOST_NETWORK default value to true
This commit changes CSI_ENABLE_HOST_NETWORK default value to true
since it was observed that cephfsvolume gets blocked when csi cephfs
nodeplugin is restarted and nodeunpublish call also hangs when
cephfs nodeplugin is using pod networking.

Updates: #8085

Signed-off-by: Rakshith R <rar@redhat.com>
2021-06-10 15:33:49 +05:30
Travis Nielsen a8764a6df1 Merge pull request #8043 from Rakshith-R/cleanup-csi-statefulset
ceph: remove obsolete statefulset functions
2021-06-03 11:06:49 -06:00
Rakshith R 35a9612041 ceph: remove obsolete statefulset functions
This commit removes obsolete statefulset create, delete,
addlabel, templateToStatefulset functions which were part of
deprecated csi support for k8s 1.13 and should have been
removed as part of https://github.com/rook/rook/pull/5982.
It also adds Unit test for templateToDeployment func.

Signed-off-by: Rakshith R <rar@redhat.com>
2021-06-03 15:36:25 +05:30
Travis Nielsen 5c332f6954 Merge pull request #8020 from Rakshith-R/csi/retry-on-k8s-err
ceph: retry starting CSI drivers on failure
2021-06-02 09:28:31 -06:00
Rakshith R 411e36900e ceph: retry starting CSI drivers on failure
CSI drivers failed to start if rook encountered error
while reading configmap with no retry.
This commit introduces retry for starting CSI driver
if there is a failure in reading configmap upto 3 times.

Fixes: #7950

Signed-off-by: Rakshith R <rar@redhat.com>
2021-06-02 19:36:55 +05:30
Rakshith R 87cf668b28 ceph: add option for adding csi pods' tolerations & affinity separately
This commit adds support for adding tolerations and affinities for
cephfs and rbd (provisioner & nodeplugin) pods separately.

Closes: #7060

Signed-off-by: Rakshith R <rar@redhat.com>
2021-06-02 11:30:38 +05:30
Travis Nielsen b0a63711f5 build: refactor to consolidate the rook.io/v1 package
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-18 19:55:37 -06:00
Rakshith R e9f0136b13 ceph: remove redundant csi statefulset template path vars
This commit removes redundant `RBDProvisionerSTSTemplatePath` and
`CephFSProvisionerSTSTemplatePath` variables which were forgetten to
be removed in https://github.com/rook/rook/pull/5982 when csi support
for k8s 1.13 was removed.

Signed-off-by: Rakshith R <rar@redhat.com>
2021-05-11 14:36:15 +05:30
Rakshith R 7987e1b8a2 ceph: update snapshot APIs from v1beta1 to v1
This commit updates external-snapshotter version to
v4.0.0 which supports snapshots v1.
Rook now defaults to enabling RBD and CephFS snapshotter
for K8s >= v1.17 and disabling it for K8s <= v1.16.
Supporting changes in documents and examples yaml files
are made.

Signed-off-by: Rakshith R <rar@redhat.com>
2021-05-05 21:04:10 +05:30
Madhu Rajanna 6bc3c1ac3a ceph: update cephcsi to v3.3.1
As we have cephcsi v3.3.1 release
with couple of bug fixes and a CVE fix
updating the cephcsi to latest release.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-04-22 13:57:01 +05:30
Rakshith R f5a3f6fa71 ceph: add option to toggle host networking in csi plugin pods
As host networking is no longer necessary for CSI driver,
it will be disabled by default.
An option CSI_ENABLE_HOST_NETWORK is added to rook-ceph-operator-config
configmap and enableCSIHostNetwork in helm chart values.yaml to
enable/disable host network in CSI plugin pods.

Closes: #7203

Signed-off-by: Rakshith R <rar@redhat.com>
2021-04-20 12:57:09 +05:30
bipuladh 7d5cc8d06b ceph: adds a container to support volume replication controller
this commit adds a new container inside of rbd-provisioner.

Signed-off-by: bipuladh <badhikar@redhat.com>
2021-04-15 21:39:26 +05:30
Madhu Rajanna fd28e48664 ceph: update csi image to v3.3.0
as we are having v3.3.0 cephcsi latest release
updating the cephcsi image to the same.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-04-15 19:36:31 +05:30
Stefan Haas 67024c1e0f ceph: upgrade ceph-csi to v3.2.1
Signed-off-by: Stefan Haas <shaas@suse.com>
2021-03-30 16:31:10 +02:00
Blaine Gardner 795124b7a8 ceph: update osds in parallel
Update OSDs in parallel per the design in
design/ceph/update-osds-in-parallel.md

The max number of OSDs updated in parallel is currently fixed at 20.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-03-29 10:55:28 -06:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Madhu Rajanna 0a81ce2251 ceph: disable CSI GRPC metrics by default
The GRPC metrics exposed by both the provisioner pod
and the node plugin pod on some port. Provisioner pod
is running on the pod network and the daemonset pods
run on the host network. sometimes starting the GRPC
metrics by default can lead to node plugin pods
crashloopback state this is due to the port conflict.
Moreover, the GRPC metrics are not for the user it's
for the one which will help to debug the time taken
by cephcsi to serve each GRPC call. Enabling it by
default won't be a good idea. So the plan is to
disable it by default, If someone faces any issue it
can be enabled later at some point in time.

closes #7378

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-03-10 11:30:50 +05:30
shenjiatong 2b63fff087 ceph: skip csi detection if CSI is disabled
Before this patch, csi detection will run whether or
not rbd and cephfs are all disabled. Run detection only
if it is necessary.

Signed-off-by: shenjiatong <yshxxsjt715@gmail.com>
2021-02-02 08:28:35 +08:00
Madhu Rajanna e305810bf0 ceph: add option to set FSGroupPolicy for CSI PVC
in kubernetes 1.19 it added a support to set FSGroupPolicy
in the csidriver object, same supported is added to both
CephFS and RBD. where user will have an option to set
the FSGroupPolicy for CephFS and RBD separately.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-01-22 14:26:56 +05:30
Madhu Rajanna 05c75b7084 ceph: add support to disable snapshotter sidecar
In some cases the user dont want to run snapshotter
container either for CephFS or RBD. In that case the
user wont install the required snapshot CRD's due
to that the snapshotter sidecar container produces
lot of noisy logs.
Snapshotter will be enabled by default for both
CephFS and RBD, but with this PR we are providing
an option to disable snapshotter sidecar deployment
either for CephFS or RBD.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-01-21 15:29:55 +05:30
Madhu Rajanna 5392bc0d07 ceph: update csi to v3.2.0
updating cephcsi to latest release version

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-12-11 20:22:36 +05:30
Madhu Rajanna ece9efd978 ceph: add mgr caps for csi-rbd-node client
If the kernel doesnt support mapping of
rbd image with deep-flatten feature,cephcsi
need to flatten the rbd image first and than
map the image on the node.

cephcsi first tries to add a task to flatten
the rbd image if its receives any permission
error it will try to call rbd CLI command which
is a blocking call.
This commit adds mgr caps to the csi-rbd-node user
so that cephcsi will add flatten task and return
immediate error to the kubelet, let kubelet retry
again,If we go with blocking rbd CLI call we may
end up having stale maps on the node in corner cases.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-12-08 17:09:18 +05:30