The ci was using a pretty old version og golangci-lint.
This updates to the latest version.
Additionally, it silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
and string format errors found by golangci-lint, while at it.
Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
enable and disable the status of rados namespace
mirroring by looking at statusCheck spec of blockpool
Signed-off-by: parth-gr <partharora1010@gmail.com>
currently if we disable the monitoring for cephblockpool
it didnt worked, as there was no check to cancel it
if the mirroring is still enabled
Add a check to disable the monitoring if mirroring is still enabled
closes: https://github.com/rook/rook/issues/14958
Signed-off-by: parth-gr <partharora1010@gmail.com>
This PR fixes an issue where disabling enableRBDStats in CephBlockPool did not remove
the pool from rbd_stats_pools, leading to unnecessary monitoring.
The update ensures that when enableRBDStats is set to false,
the pool is properly removed from rbd_stats_pools, optimizing resource tracking.
Signed-off-by: Oded Viner <oviner@redhat.com>
This is similar to #14052 we did for radosnamespace
and this is an extension to support cleanup
at the blockpool level to cleanup the images
and the snapshots in a pool.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Modify the CR to allow mirroring of an rados namespace
to a differently named namespace on the remote cluster
1) enable rados namesapce mirroring only
if the blockpool mirrroing is enabled
2) disable blockpool mirroing only if
all the namesapce mirroing is disabled
if the rbd mirroring fails and ceph version is not supported
provide a error message with supported version details
and reason of failing
Signed-off-by: parth-gr <partharora1010@gmail.com>
default application name is updated inside the `CreatePool` method. Send
pool spec as address in order to preserve this change.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Updating the device class swallowed any error if updated
for the pool. The error was not even logged, so we couldn't
troubleshoot why the new crush rule was not applied.
Log the error for troubleshooting and also fail the pool
reconcile since the desired configuration was not applied.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
make the kubectl output for crds more verbose.
Following changes was done for each crd:
CephBlockPool:
Added print columns:
- Type
- FailureDomain
- Replication
- EC-CodingChunks
- EC-DataChunks
- Age
Extended info field of status to include:
type and failureDomain
CephObjectStore:
Added print columns:
- Endpoint
- SecureEndpoint
- Age
Added Age to print columns of following crds:
CephObjectStoreUser, CephObjectZoneGroup,
CephObjectZone, CephBucketTopic, CephClient,
CephRBDMirror, CephFilesystemMirror,
CephFilesystemSubVolumeGroup
Signed-off-by: NymanRobin <nyman.robin@gmail.com>
Pool name either come from CephBlockPool name or
form the CephBlockPool Spec. Use the correct name
when trying to disable mirroring.
Signed-off-by: sp98 <sapillai@redhat.com>
For image mode mirroring, if cephBlockPool.Pool.Spec.Mirroring.Enable
is set to false, then remove the peer cluster and disable mirroring on
all the pool if the user has disabled mirroring on all the pool images.
If mirroring is not disabled on all the pool images, then reconcile will
fail asking the users to manually disable mirroring on those images.
Signed-off-by: sp98 <sapillai@redhat.com>
Rook has been setting the application automatically on all
pools to rbd for CephBlockPools, rook-ceph-rgw for
CephObjectStores, mgr on the built-in .mgr pool,
and nfs on the built-in .nfs pool.
The legacy pool device_health_metrics is long gone
from Pacific which is no longer supported, so we can
remove special handling for that pool in the upgrade
guide and in the code.
The application setting is now available on the pool spec
although it is not expected to commonly need to override
the default applications set by Rook.
The application for CephFilesystem pools is now being
set to cephfs, where it was previously blank.
Signed-off-by: travisn <tnielsen@redhat.com>
Rook currently knows rbd pools only if they defined through CephBlockPool, so it overrides the config
mgr/prometheus/rbd_stats_pools which may have already been set for pools external to Rook. Check config
before running set command and if there some existing pools append them to enableStatsForCephBlockPools.
Signed-off-by: Avan Thakkar <athakkar@redhat.com>
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor
Signed-off-by: parth-gr <paarora@redhat.com>
`rbd pool init` cmd initializes pool for rbd images.
This commit makes modification to init only
rbd application pools, since it is not required
by other pools like ".nfs",".mgr" & "mgr_devicehealth".
Signed-off-by: Rakshith R <rar@redhat.com>
During deletion of a CephBlockPool CR, the underlying ceph
pool was not being deleted. During the deletion sequence
this is due to the spec being refreshed upon a call to
check for dependents in ReportDeletionNotBlockedDueToDependents().
The pool name was then coming up blank during the deletion
request and rook of course then wasn't finding the pool
to delete.
Now Rook properly sets the pool name in the named spec
every time it is converted internally from a CephBlockPool
type to a NamedPoolSpec.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CSI package needs to load clusterInfo, today this code is in the mon
package which makes the call of LoadClusterInfo impossible without
having a circular import.
Signed-off-by: Sébastien Han <seb@redhat.com>
This introduces a new CRD to add the ability
to create rados namespace for a given
ceph block pool. Typically the name of the pool
is the name of the blockpool created by rook.
Closes: #7035
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
In Quincy, the device_health_metrics pool is renamed to .mgr.
In the example cluster-test.yaml, we therefore update the
example to create the pool with the new name.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller
Closes: https://github.com/rook/rook/issues/9673
Signed-off-by: parth-gr <paarora@redhat.com>
The reconcile was skipping updating most pool properties for
object stores. The implementation of pools between the file,
object, and pool controllers had some duplicate code, so
this change also factors out the common code for better
reuse in a single place. Anytime a pool is created or updated,
it will now consistently update all the pool properties
that are expected to be modifiable.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The built-in pools device_health_metrics and .nfs created by ceph
need to be configured for replicas, failure domain, etc.
To support this, we allow the pool to be created as a CR.
Since K8s does not support underscores in the resource names
the operator must translate this special pool name into
the name expected by ceph.
This also sets the basis for allowing filesystem data
pools to specify the desired pool name instead of requiring
a generated name.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When CephBlockPool, CephFilesystem, or CephObjectStore resources are
deleted after removing their finalizer, the code path to stop monitoring
was not stopping monitoring since a non-present resource does not have a
name and namespace attached. When the object is deleted, ensure the
internal representation used to stop monitoring has a name and namespace
to fix the issue.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Clean up the code used to stop health checkers for all controllers
(pool, file, object). Health checkers should now be stopped when
removing the finalizer for a forced deletion when the CephCluster does
not exist. This prevents leaking a running health checker for a resource
that is going to be imminently removed.
Also tidy the health checker stopping code so that it is similar for all
3 controllers. Of note, the object controller now uses namespace and
name for the object health checker, which would create a problem for
users who create a CephObjectStore with the same name in different
namespaces.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This is done in order to prevent deadlock when parallel
PVC create requests are issued on a new uninitialized
rbd block pool due to https://tracker.ceph.com/issues/52537.
Fixes: #8696
Signed-off-by: Rakshith R <rar@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
The ceph-block-pool controller was missing the network spec from the
cephcluster that contains all the details about networking.
So when the pool was deleting on multus the controller was not proxying
the rbd command to the mgr pod but executed the command in the operator
pod which does not have the network annotation and then connect to the
ceph cluster.
Signed-off-by: Sébastien Han <seb@redhat.com>