This commit introduces a new configuration option for
ceph csi driver to enable hostpath mounting of /etc/selinux
directory from the cluster node where csi plugin pods are
running, which inturn help the csi driver to specify
selinux-related mount options like context.
Ref# https://github.com/ceph/ceph-csi/issues/2295
The default value for this configuration is true and if cluster
nodes are running without selinux enabled, an admin can deploy
csi pods by specifying this option to `false` which skip the
host path mounting for the csi pods.
Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil daemonset functions. By
this, we can handle cancellation during API call of daemonset resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit makes sure csi provisioner node-affinity is applied
to the csi version check job as well, similar to the how
the existing csi provisioner toleration is applied to the job.
Fixes: #8323
Signed-off-by: Rakshith R <rar@redhat.com>
The CSI controller loop watches ConfigMaps and CephClusters. When a
CephCluster triggers the reconcile loop, the Get call for the ConfigMap
will fail since the name in the request object is that of the
CephCluster. Setting the name and namespace to the constants for the
ConfigMap name and the operator namespace resolves this.
Closes: https://github.com/rook/rook/issues/8958
Signed-off-by: Renan Campos <rcampos@redhat.com>
In Rook v1.8 the min version of K8s supported is updated to 1.16.
Users running on older versions of K8s are recommended to update
to 1.16 or newer before updating to Rook v1.8.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
As like RBD, CephFS provisioner pod need not to
run as privileged. as its not doing any operation
like plugin pods which does mounting and unmounting
removing the permissions for the same.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
added new option to set the provisioner replicas.
with this new option the user/admin can choose
how many replicas he want for provisioner pod if
number of nodes is greater than 1.
fixes#8153
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
With the release of Ceph CSI 3.4.0, Ceph CSI project came up
with a new support policy where we support only versions >= 3.3.0
The supported window of Ceph CSI versions is known as "N.(x-1)":
(N (Latest major release) . (x (Latest minor release) - 1)).
For example, if Ceph CSI latest major version is 3.4.0 today,
support is provided for the versions above 3.3.0.
If users are running an unsupported Ceph CSI version, they will be
asked to upgrade when requesting support for the cluster.
This PR lift the minimum supported version of Ceph CSI to 3.3.0
in this repo.
Fix https://github.com/rook/rook/issues/8709
Ref #
https://github.com/ceph/ceph-csi/releases/tag/v3.4.0https://github.com/ceph/ceph-csi/#known-to-work-co-platforms
Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
We don't need to use github.com/ghodss/yaml since
"k8s.io/apimachinery/pkg/util/yaml" provides the same functionality and
we already import it.
Signed-off-by: Sébastien Han <seb@redhat.com>
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters
This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Thanks to Golang 1.16, we can now embed files in the Go binary. This
means we don't need to add the CSI templates files to the container
image. They are added in the Go binary at build time.
The existing location of the template must be in the package calling it.
So they moved to pkg/operator/ceph/csi/template. All the files have been
symlinked back to cluster/examples/kubernetes/ceph/csi/template.
Closes: https://github.com/rook/rook/issues/7609
Signed-off-by: Sébastien Han <seb@redhat.com>
log the successful message about
starting CSIDriver after creating
the daemonset and deployment objects.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
The `None` value Indicates that volumes will be mounted
with no modifications, as the CSI volume driver does not
support these operations. While volumes are provisioned
by the CephFS CSI driver the global permissions are set
on the volume and we dont expect the Fsgroup policy or
check from CO side to play a role here.
Ref #ceph/ceph-csi/../internal/cephfs/nodeserver.go#L190
```
!csicommon.MountOptionContains(fuseMountOptions, readOnly) {
// #nosec - allow anyone to write inside the stagingtarget path
err = os.Chmod(stagingTargetPath, 0o777)
```
The current default value ie `ReadWriteOnceWithFSType` cause
volumes to be examined to determine if volume ownership and permissions
should be modified to match the pod's security policy. Changes could occur
if the fsType is defined and the persistent volume's accessModes contains
ReadWriteOnce.
Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
This PR updates the required RBAC, templates,
CSI image version and examples for new
cephcsi v3.4.0 release.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Update the CSIDriver object to v1 from betav1 as it will be
deprecated in v1.19+ and become unavailable in v1.22+
Signed-off-by: Yug <yuggupta27@gmail.com>
This commit moves SetParams and ValidateCSIParam func
into go routine/lock, since it modifies/reads global CSIParam
variable.
Signed-off-by: Rakshith R <rar@redhat.com>
This commit changes CSI_ENABLE_HOST_NETWORK default value to true
since it was observed that cephfsvolume gets blocked when csi cephfs
nodeplugin is restarted and nodeunpublish call also hangs when
cephfs nodeplugin is using pod networking.
Updates: #8085
Signed-off-by: Rakshith R <rar@redhat.com>
This commit removes obsolete statefulset create, delete,
addlabel, templateToStatefulset functions which were part of
deprecated csi support for k8s 1.13 and should have been
removed as part of https://github.com/rook/rook/pull/5982.
It also adds Unit test for templateToDeployment func.
Signed-off-by: Rakshith R <rar@redhat.com>
CSI drivers failed to start if rook encountered error
while reading configmap with no retry.
This commit introduces retry for starting CSI driver
if there is a failure in reading configmap upto 3 times.
Fixes: #7950
Signed-off-by: Rakshith R <rar@redhat.com>
This commit adds support for adding tolerations and affinities for
cephfs and rbd (provisioner & nodeplugin) pods separately.
Closes: #7060
Signed-off-by: Rakshith R <rar@redhat.com>
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit removes redundant `RBDProvisionerSTSTemplatePath` and
`CephFSProvisionerSTSTemplatePath` variables which were forgetten to
be removed in https://github.com/rook/rook/pull/5982 when csi support
for k8s 1.13 was removed.
Signed-off-by: Rakshith R <rar@redhat.com>
This commit updates external-snapshotter version to
v4.0.0 which supports snapshots v1.
Rook now defaults to enabling RBD and CephFS snapshotter
for K8s >= v1.17 and disabling it for K8s <= v1.16.
Supporting changes in documents and examples yaml files
are made.
Signed-off-by: Rakshith R <rar@redhat.com>
As we have cephcsi v3.3.1 release
with couple of bug fixes and a CVE fix
updating the cephcsi to latest release.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
As host networking is no longer necessary for CSI driver,
it will be disabled by default.
An option CSI_ENABLE_HOST_NETWORK is added to rook-ceph-operator-config
configmap and enableCSIHostNetwork in helm chart values.yaml to
enable/disable host network in CSI plugin pods.
Closes: #7203
Signed-off-by: Rakshith R <rar@redhat.com>
Update OSDs in parallel per the design in
design/ceph/update-osds-in-parallel.md
The max number of OSDs updated in parallel is currently fixed at 20.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The GRPC metrics exposed by both the provisioner pod
and the node plugin pod on some port. Provisioner pod
is running on the pod network and the daemonset pods
run on the host network. sometimes starting the GRPC
metrics by default can lead to node plugin pods
crashloopback state this is due to the port conflict.
Moreover, the GRPC metrics are not for the user it's
for the one which will help to debug the time taken
by cephcsi to serve each GRPC call. Enabling it by
default won't be a good idea. So the plan is to
disable it by default, If someone faces any issue it
can be enabled later at some point in time.
closes#7378
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Before this patch, csi detection will run whether or
not rbd and cephfs are all disabled. Run detection only
if it is necessary.
Signed-off-by: shenjiatong <yshxxsjt715@gmail.com>
in kubernetes 1.19 it added a support to set FSGroupPolicy
in the csidriver object, same supported is added to both
CephFS and RBD. where user will have an option to set
the FSGroupPolicy for CephFS and RBD separately.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
In some cases the user dont want to run snapshotter
container either for CephFS or RBD. In that case the
user wont install the required snapshot CRD's due
to that the snapshotter sidecar container produces
lot of noisy logs.
Snapshotter will be enabled by default for both
CephFS and RBD, but with this PR we are providing
an option to disable snapshotter sidecar deployment
either for CephFS or RBD.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
If the kernel doesnt support mapping of
rbd image with deep-flatten feature,cephcsi
need to flatten the rbd image first and than
map the image on the node.
cephcsi first tries to add a task to flatten
the rbd image if its receives any permission
error it will try to call rbd CLI command which
is a blocking call.
This commit adds mgr caps to the csi-rbd-node user
so that cephcsi will add flatten task and return
immediate error to the kubelet, let kubelet retry
again,If we go with blocking rbd CLI call we may
end up having stale maps on the node in corner cases.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>