having csi-addons enabled by default is causing random
pod restart on non-openshift cluster. Let's disable
it by default.
Signed-off-by: subhamkrai <srai@redhat.com>
this command adds some examples on how users can add/update
the settings based on new way of managing CSI resources.
Signed-off-by: subhamkrai <srai@redhat.com>
The ROOK_RECONCILE_CONCURRENT_CLUSTERS feature was implemented
in v1.19. This feature has been stable, with no related issues
reported. The feature is tested in the CI with no known
stability issues. Let's declare this feature as stable.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Going forward, admin will manage the csi operator
CR's and rook will only manage Ceph Connection cr
and client Profile cr.
The old csi driver is completely removed from Rook
and can no longer be used starting in Rook v1.20.
The upgrade guide will contain the needed transition steps
for managing the csi operator settings.
Signed-off-by: subhamkrai <srai@redhat.com>
Clean up stale CRUSH rules after the Ceph mgr starts so rules left behind by pool failure-domain or device-class changes do not accumulate indefinitely.
The cleanup is guarded by a package-level RWMutex. Pool create and update paths hold the read lock while creating and assigning CRUSH rules, while cluster-wide cleanup holds the write lock before listing pools and deleting unused rules. This keeps pool reconciles parallel with each other while preventing cleanup from deleting a rule that another reconcile has just created but not yet attached to a pool.
Keep direct pool-delete cleanup for the pool's current CRUSH rule, make the cluster-wide cleanup best-effort across all unused rules, and add an operator-level ROOK_DELETE_UNUSED_CRUSH_RULES setting for clusters that need to leave unused custom rules in place.
Document the operator and Helm settings, regenerate the Helm chart docs, and add a pending release note for the default cleanup behavior.
Signed-off-by: Asish Kumar <officialasishkumar@gmail.com>
Updated the following csi sidecars to their latest available versions:
- csi-attacher: v4.11.0
- csi-snapshotter: v8.5.0
- csi-resizer: v2.1.0
- csi-provisioner: v6.1.1
- csi-node-driver-registrar: v2.16.0
Signed-off-by: Praveen M <m.praveen@ibm.com>
since node fencing is disabled in rook for sometime
and this feature is implemented in csi. So, this
commits remove obsolete code related to node loss
Signed-off-by: subhamkrai <srai@redhat.com>
The cephcluster controller allows setting concurrent reconciles
with a setting in the operator configmap.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
this commits enable the csi-operator by default,
moving it from experimental to stable.
Also, disable the csi-operator chart from generating
rbac in common.yaml
Signed-off-by: subhamkrai <srai@redhat.com>
This change addresses a permission issue where mon pods crashloop
on some Kubernetes setups with SELinux enabled, even when
ROOK_HOSTPATH_REQUIRES_PRIVILEGED is set.
To resolve this, a new function makeMonSecurityContext() was introduced
to explicitly set runAsUser: 0 at the pod level for mon pods
when the environment variable ROOK_CEPH_MON_RUN_AS_ROOT is set to true.
Additionally, the function previously named PodSecurityContext() was renamed
to DefaultContainerSecurityContext() to avoid confusion between
container-level and pod-level security context configuration. All container
SecurityContext usages across Ceph daemons were updated to reflect this change.
This ensures the root user configuration is applied
automatically and consistently in environments where it is required
for mon pod startup, while preserving clear separation of container
and pod-level security logic.
Signed-off-by: Patryk Rostkowski <patrostkowski@gmail.com>
The Kubernetes CSI sidecars have had several releases that were not
included in deployments by Rook yet, update them to the versions that
are available today:
- csi-attacher:v4.8.1
- csi-provisioner:v5.2.0
- csi-resizer:v1.13.2
- csi-snapshotter:v8.2.1
This change is important, because Ceph-CSI will implement the new
Controller.GetSnapshot CSI procedure. A bug in csi-lib-utils causes a
panic when a ControllerCapability is provided, but not (yet) known to
the CSI sidecars. The updated sidecars consume a version of
csi-lib-utils with a fix for that panic.
See-also: kubernetes-csi/csi-lib-utils#188
Signed-off-by: Niels de Vos <ndevos@ibm.com>
When host network is enabled, the operator needs to set the
dns policy to ClusterFirstWithHostNet so the request to the
rgw endpoint will resolve properly.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator settings loaded from the configmap have proven
inefficient for load time and frequently checking the configmap.
To avoid this ineffenciency, the configmap is only loaded once
each time it is created or updated. The values in the configmap
are applied as environment variables, which then are very efficient
to query throughout the various controllers, without needing
to be concerned about loading the configmap again.
Co-authored-by: Dmitry Mishin <dmitry.mishin@gmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Implement an allow list mechanism that disables potentially unsafe OBC
fields by default. OBC fields beyond `maxObjects` and `maxSize` don't
neatly fit into the OBC framework as it was originally envisioned and
implemented.
Some of the newly added configs could allow users to cause confusion for
themselves. Others might allow users to hijack others buckets. Some
might allow bricking the entire S3 store.
Out of an abundance of safety, allow-list the known-safe options by
default, and require administrators to enable potentially troublesome
options via the new operator-level config
`ROOK_OBC_ALLOW_ADDITIONAL_CONFIG_FIELDS`.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The Kubernetes CSI sidecars have had several releases that were not
included in deployments by Rook yet, update them to the versions that
are available today:
- csi-node-driver-registrar:v2.13.0
- csi-provisioner:v5.1.0
- csi-attacher:v4.8.0
- csi-resizer:v1.13.1
Signed-off-by: Niels de Vos <ndevos@ibm.com>
cephcsi fixed a bug related to data loss
and its fixed in 3.12.3 release, This commit
updates the cephcsi to 3.12.3 release.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
The version checks for the csi driver are removed now
since they are all obsolete. The K8s version and cephcsi
versions are no longer checked. Anyway, the move to the
csi operator would take ownership of version checks
needed in the future, so for now we simplify rook
deployment of the csi driver.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Finish the process of deprecating holder pods by removing Rook's ability
to deploy them. The intent of this change is to make the most
superficial changes possible to accomplish this. There are still
remnants of code in Rook (particularly the CSI controller) that helped
configure or deploy holder pods. Due to the risk of breaking some
features, cleanup work of hose remnants will be deferred for future
work.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The ROOK_ENFORCE_HOST_NETWORK option was implemented recently
and now we add the helm setting to expose this new setting
in the rook chart.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This adds an operator config setting ROOK_REVISION_HISTORY_LIMIT
defaulting to kubernetes'value for RevisionHistoryLimit.
If configured, the provided value will be used as RevisionHistoryLimit
for all Deployments rook creates.
Fixes: #12722
Signed-off-by: Michael Adam <obnox@samba.org>
Alerting on controller-runtime's workqueue_depth can be useful for
debugging controllers. Also having a prometheus target for a pod gives
another data point that the system is working as expected. It is useful
for uptime alerts.
Make the bind address configurable via the configmap while still retaining the default
behavior that it is disabled.
Resolves: #14538
Signed-off-by: Justin Cichra <jcichra@cloudflare.com>