Commit Graph
361 Commits
Author SHA1 Message Date
Travis Nielsen e068e156ce osd: enable osd ok-to-stop checks on single node for three osds
If there are at least three OSDs on a single node, we should
treat it as a potential production cluster and perform
the ok-to-stop checks during reconcile. Otherwise,
it may cause instability during upgrades on
single-node clusters.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2025-02-05 12:08:48 -07:00
Travis Nielsen 16daa7dc44 Merge pull request #15184 from OdedViner/err_msg_ok_to_stop
operator: improve operator error logging for ok-to-stop failures
2025-01-24 08:13:20 -07:00
Eng Zer Jun e68a7ea561 build: replace golang.org/x/exp/slices with stdlib slices
The experimental functions in `golang.org/x/exp/slices` are now
available in the standard library in Go 1.21.

Reference: https://go.dev/doc/go1.21#slices
Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
2025-01-21 15:02:35 +08:00
Oded Viner 3d7f5cff65 operator: improve operator error logging for ok-to-stop failures
this PR improves the error logging when ok-to-stop requests fail
in the Rook operator. Instead of only showing a generic exit status

Signed-off-by: Oded Viner <oviner@redhat.com>
2025-01-19 14:46:57 +02:00
Michael Adam 26b4fb06d5 core: add function roundupSizeMiB
This change  refactors inline roundup code to a function
roundup_size_MiB that takes a size(in bytes) as argument and returns
the smallest number of Mibibytes larger than or equal to the given size.

Signed-off-by: Michael Adam <obnox@samba.org>
2024-12-16 15:30:47 +01:00
df511fb58f ci: update golangci-lint to the latest version (v1.62)
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.

Additionally, it  silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
 and string format errors found by golangci-lint, while at it.

Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
2024-12-14 14:47:30 +01:00
parth-gr 996a88f902 rbdmirror: rbd import cmd stuck forever
if the import from the peer cluster doesnt
work, because of some reasons like
cluster health, network, etc

the import cmd stuck forever and which
forever stucks the blockppol reconcile

Signed-off-by: parth-gr <partharora1010@gmail.com>
2024-11-25 17:04:11 +05:30
parth-gr ed5013e9c6 rbdmirror: add disable check for status check
enable and disable the status of rados namespace
mirroring by looking at statusCheck spec of blockpool

Signed-off-by: parth-gr <partharora1010@gmail.com>
2024-11-06 22:42:46 +05:30
Travis Nielsen 7692cae359 Merge pull request #14896 from parth-gr/rbd-mirror-rados-monitor
rbd: enable periodic monitoring for rados namespace mirroring
2024-11-05 09:15:32 -07:00
parth-gr 9785ef5c8d rbdmirror: enable periodic monitoring for rados namespace
enable monitoring for rados namespace

Signed-off-by: parth-gr <partharora1010@gmail.com>
2024-11-05 13:45:02 +05:30
Travis Nielsen 99d60c5bba mds: wait for mds standby upgrade for same fs
The upgrade of the mds daemons pauses to wait for the
stopping of standby daemons before the filesystem upgrade
is completed. If there are multiple filesystems, there
could be standbys for other filesystems that will not be
stopped at the time of the upgrade of another filesystem.
Thus, the standbys may not upgrade and be stuck on the
previous version. Now, the standby upgrade will only wait
for standbys for the same filesystem.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-11-01 14:03:13 -06:00
Madhu Rajanna 4e29717317 core: cleanup blockpool with annotation
This is similar to #14052 we did for radosnamespace
and this is an extension to support cleanup
at the blockpool level to cleanup the images
and the snapshots in a pool.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2024-10-23 12:20:31 +02:00
Travis Nielsen 0fa21969df Merge pull request #14701 from parth-gr/rbd-mirror-rados
rbdmirror: enable rbd rados namespace mirroring
2024-10-10 09:35:22 -06:00
parth-gr 2cef47a9b7 rbdmirror: enable rbd rados namespace mirroring
Modify the CR to allow mirroring of an rados namespace
to a differently named namespace on the remote cluster

1) enable rados namesapce mirroring only
if the blockpool mirrroing is enabled

2) disable blockpool mirroing only if
all the namesapce mirroing is disabled

if the rbd mirroring fails and ceph version is not supported
provide a error message with supported version details
and reason of failing

Signed-off-by: parth-gr <partharora1010@gmail.com>
2024-10-10 14:23:11 +05:30
Travis Nielsen b665d7a7b7 core: remove support for ceph quincy
Given that Ceph Quincy (v17) is past end of life,
remove Quincy from the supported Ceph versions,
examples, and documentation.

Supported versions now include only Reef and Squid.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-10-03 11:12:55 -06:00
Santosh Pillai 1a95d6f533 core: preserve pool application name change
default application name is updated inside the `CreatePool` method. Send
pool spec as address in order to preserve this change.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2024-09-25 12:12:01 +05:30
Travis Nielsen b9188c793a pool: allow negative step num in crush rule
The crush rules may have a negative step num.
Rook had assumed negative values were not possible,
but just had not been encountered previously in
a custom crush rule.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-09-09 13:27:54 -06:00
Blaine Gardner afdfffac67 Merge pull request #14435 from parth-gr/osd-resize
osd: reweight osd while resizing
2024-08-20 12:14:22 -06:00
parth-gr 17cfda5bd6 osd: reweight osd while resizing
osd got resized by cryptsetup bluestore cmd
but it should also be reweight to balance the pgs properly

closes: https://github.com/rook/rook/issues/14430

Signed-off-by: parth-gr <partharora1010@gmail.com>
2024-08-20 22:52:33 +05:30
Travis Nielsen d91b8073c5 Merge pull request #14488 from sp98/set-min-compat-client-to-reef
core: set min-compat-client to reef for upmap-read
2024-07-29 11:28:29 -06:00
sp98 ee5e710912 core: set min-compat-client to reef for upmap-read
If upmap-read balancer mode is required, then set
min-compat-client to reef

Signed-off-by: sp98 <sapillai@redhat.com>
2024-07-29 22:03:37 +05:30
Travis Nielsen 88e952a3ac osd: update the device class if desired state changes
Normally the device class of an OSD is determined at provisioning
and is not updated thereafter. In some scenarios the admin may
want to force update the device class to a new value. The device
classes can be updated by first setting allowDeviceClassUpdate
in the storage spec of the cephcluster, then updating the
device class specified on the deviceSets or other OSDs.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-07-17 13:54:21 -06:00
Travis Nielsen 0bea31ef7f pool: return error if device class update fails
Updating the device class swallowed any error if updated
for the pool. The error was not even logged, so we couldn't
troubleshoot why the new crush rule was not applied.
Log the error for troubleshooting and also fail the pool
reconcile since the desired configuration was not applied.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-07-10 12:59:43 -06:00
Travis Nielsen b92bbba671 pool: skip updating crush rules for stretch clusters
Pools in stretch clusters must all specify the same
CRUSH rule. No pools can use a different rule. When there
is a change in the device class, we do not even expect to update
the crush rules in a stretch cluster. Different device classes
are not supported in stretch clusters, and it's expected to be
a homogenous environment. Therefore, skip all crush rule updates
in stretch clusters.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-07-10 11:08:31 -06:00
Blaine Gardner 4b51c2c3bd Merge pull request #14322 from ideepika/wip-disablecrush-updates
enableCrushUpdates option: rook operator would not be able to update crush rules by default if not set
2024-06-17 12:20:26 -06:00
Deepika Upadhyay 41116be919 pool: get the exact deviceClass name instead of crushroot+deviceClass
this lead to creation of 2 crush rules because of difference in
deviceClass names:
eg:
creating a new crush rule for changed deviceClass
("default~hdd"-->"hdd") on crush rule "test-crush-bug-az-ab-4"

Signed-off-by: Deepika Upadhyay <deepikaupadhyay01@gmail.com>
2024-06-13 01:03:20 +05:30
Deepika Upadhyay d71f9c238f pool: add option to enableCrushUpdates to pool if required
this will be disabled by default but in scenarios where the user want to
update failureDomain, DeviceClass etc, this option can be enabled, to be
noted this can lead to lot of data rebalancing and remapping. Use with
caution

Signed-off-by: Deepika Upadhyay <deepika.upadhyay@clyso.com>
2024-06-12 22:48:25 +05:30
Travis Nielsen e479b051bc osd: configure cluster full settings when osds fill up
When the clusters reach full, nearfull, or backfill full thresholds
ceph will raise health warnings and stop allowing IO or backfill
depending on the threshold. These settings require special ceph
commands instead of being generic ceph config. Allow these settings
to be set from the CephCluster CR in the spec.storage
section.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-05-30 12:10:41 -06:00
Satoru Takeuchi f776c0b35e core: remove pacific specific code
Since pacific is no longer supported, we can remove pacific specific
code.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2024-05-15 02:26:33 +00:00
Travis Nielsen 61332caa4b Merge pull request #14049 from NymanRobin/crd-output-enhancements
operator: make CephBlockPoolRadosNamespace and CephFilesystemSubVolumeGroup get outputs more verbose
2024-04-16 12:00:15 -06:00
NymanRobin 70a7e61da5 operator: make crds outputs more verbose
The following printcolumns were added

CephBlockPoolRadosNamespace: Phase,
BlockPoolName, Age
CephFilesystemSubVolumeGroup: Phase,
FilesystemName, PinningConfig, Age

Also improved the error messages in
SubVolumeGroupClient

Signed-off-by: NymanRobin <nyman.robin@gmail.com>
2024-04-16 08:40:16 +03:00
Travis Nielsen 499a09ed72 Merge pull request #14052 from sp98/cleanup-radosnamespace
Cleanup RADOS namespace with forced deletion annotation
2024-04-12 11:11:47 -06:00
sp98 f6b1449faa core: cephblockpoolRadosNamespace cleanup
Clean up pool images and snapshots in the
radosnamespace

Signed-off-by: sp98 <sapillai@redhat.com>
2024-04-12 21:49:49 +05:30
Travis Nielsen 309b164c49 test: reduce wait time for mds standby test
The mds standby test was timing out after six seconds on
three different attempts, causing a single unit test to
take almost 20 seconds. Now we reduce the wait time to
a total of three seconds.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-04-11 16:19:39 -06:00
Madhu Rajanna f57a8b6bbe subvolumegroup: add support for size and datapool
cephfs subvolumegroup supports creating svg
with quota and the datapool, This PR adds the
support for the same.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2024-04-10 16:40:56 +02:00
sp98 ef00fdac53 core: subvolumegroup clean up
Cleanup the resources created by subvolumegroup
when its deleted. Following resources will be cleaned up:
- OMAP value
- OMAP keys
- Clones
- Snapshots
- Subvolumes

Signed-off-by: sp98 <sapillai@redhat.com>
2024-04-10 17:34:45 +05:30
sp98 a37c47641f core: disable mirroring in image mode
For image mode mirroring, if cephBlockPool.Pool.Spec.Mirroring.Enable
is set to false, then remove the peer cluster and disable mirroring on
all the pool if the user has disabled mirroring on all the pool images.

If mirroring is not disabled on all the pool images, then reconcile will
fail asking the users to manually disable mirroring on those images.

Signed-off-by: sp98 <sapillai@redhat.com>
2024-03-19 14:20:51 +05:30
travisn 90bec31203 pool: skip crush rule update when not needed
The crush rule will be updated when the failure domain changes.
This update can be skipped when the expected crush rule is
already configured. Otherwise, the log has a confusing message
that makes it appear that the crush rule was updated
when in fact it was not.

Signed-off-by: travisn <tnielsen@redhat.com>
2024-02-16 12:31:38 -07:00
travisn 806608cdbc pool: allow setting the application on a pool
Rook has been setting the application automatically on all
pools to rbd for CephBlockPools, rook-ceph-rgw for
CephObjectStores, mgr on the built-in .mgr pool,
and nfs on the built-in .nfs pool.

The legacy pool device_health_metrics is long gone
from Pacific which is no longer supported, so we can
remove special handling for that pool in the upgrade
guide and in the code.

The application setting is now available on the pool spec
although it is not expected to commonly need to override
the default applications set by Rook.

The application for CephFilesystem pools is now being
set to cephfs, where it was previously blank.

Signed-off-by: travisn <tnielsen@redhat.com>
2024-02-14 16:53:25 -07:00
parth-gr c9fd382c13 subvolumegroup: add pinning spec in subvolumegroup CRD
subvolumegroup can be pinned by pintype and pinsetting
So adding the spec to enhance its functionality

Closes: https://github.com/rook/rook/issues/12607

Signed-off-by: parth-gr <paarora@redhat.com>
2023-12-11 21:54:25 +05:30
Alexander Trost 3097455d78 Merge pull request #13246 from koor-tech/ceph_config_via_cluster_crd_impl
operator: allow setting ceph config options via ceph cluster crd
2023-12-02 11:27:20 +01:00
Alexander Trost 4ed35d6bd6 operator: allow setting ceph config options via ceph cluster crd
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2023-12-02 00:05:48 +01:00
parth-gr 939a1b7d20 subvolumegroup: add name spec in subvolumegroup
Originally we create it using this cmd
ceph fs subvolume create <vol_name> <subvol_name>
So we can have 2 variables filesystem and subvolume name,
Currently the CR doesn't allow us to make subvolume-name
as constant as needed to "csi" because of k8s limitations

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-30 20:39:11 +05:30
Rakshith R 59cb0dd4bf csi: add CSIDriverOptions section in cephCluster CR
This commit adds new CSIDriverOptions section in
cephCluster CR. This section contains settings
for read affinity and kernel+fuse Mount options
These settings will be injected directly into
rook-ceph-csi-config cm to be applicable per
ceph cluster.

Signed-off-by: Rakshith R <rar@redhat.com>
2023-11-29 19:34:28 +05:30
Ryotaro Banno 0ad88e9d68 core: add pgHealthyRegex to DisruptionManagementSpec
This patch adds `pgHealthyRegex` field to DisruptionManagementSpec.
`pgHealthyRegex` is a regular expression that is used to determine which
PG states should be considered healthy. The default value of
`pgHealthyRegex` is:

        ^(active\+clean|active\+clean\+scrubbing|active\+clean\+scrubbing\+deep)$

which is effectively the same as before.

Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
2023-11-22 00:19:04 +00:00
travisn 03d077aa6b core: remove support for ceph pacific
Pacific is end of life and no longer necessary to
support in Rook with v1.13.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-11-14 17:07:03 -07:00
subhamkrai 133154d192 pool: allow updating deviceClass on existing pool
This commits add check to update the deviceClass on
existing pool.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-11-13 20:40:21 +05:30
travisn f7094d1d13 file: disable active standby when set to false
The activeStandby property of the filesystem CR was not
taking effect when changed from true to false. The standby
was only being enabled when true, but never applied when
changed to false.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-10-02 08:27:14 -06:00
Travis Nielsen 0ae3800929 Merge pull request #12770 from sp98/osd-migration-part2
osd: replace existing OSDs to use new store
2023-09-05 08:57:20 -06:00
sp98 11b8d10a5a osd: replace existing OSDs to use new store
This follow up the #12507
- fixes replacing of encrypted OSDs.
- Updates OSD status at the end of reconcile

Signed-off-by: sp98 <sapillai@redhat.com>
2023-09-01 21:07:38 +05:30