few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
The requireMsgr2 setting is not available in 1.10, therefore
the setting must be removed from the cephcluster CR
when the 1.10 cluster is created before the upgrade.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Enabling msgr v2 and disabling msgr v1 currently requires enabling
either encryption on the wire or compression on the wire.
As more clients are running on the latest kernel, allow
the clients to run on v2 even when encryption and compression
are not enabled. Clusters that are fully running on v2
will more easily be able to change configuration between
enabling or disabling msgr v2 features.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the v1.11 release approaching, the upgrade tests in
master are not updated to upgrade starting from v1.10
instead of from v1.9.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The machine disruption budgets for handling openshift
machines and machinesets are now removed since they
have been unused and unmaintained since implemented.
This feature is expected to be handled with the more
common Pod Disruption Budgets. A workaround is for the
cluster admin to set up their machine sets so they
match the zone topology. See the original design
doc from the feature here:
https://github.com/rook/rook/blob/master/design/ceph/ceph-openshift-fencing-mitigation.md
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
will check if the podrestartcount is greater than 1,
If it is we will alert it and fail the CI
It is important to understand intermittent failures
to avoid too many false positives
Closes: https://github.com/rook/rook/issues/11380
Signed-off-by: parth-gr <paarora@redhat.com>
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md
Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For the upgrade case, Rook can fail to set status information
on CephObjectStores after CRDs are updated but before the operator is
updated. To fix this, merely allow the slices to be null in the
CephObjectStore's status.endpoints.
This PR seems to be aggravating the helm filesystem upgrade test. Allow
30 more seconds for CRDs to be done before bailing. In debugging, the
filesystem often needed only an extra 3 seconds to succeed.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
enabling logCollector by default will enable the coredump
generation in case a process terminates with a segmentation fault.
Signed-off-by: gauravsitlani <gauravsitlani@riseup.net>
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.
Full discussion: https://github.com/rook/rook/issues/11031
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
ci is more resource intensive.
Signed-off-by: Rakshith R <rar@redhat.com>
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
The volume replication operator is being moved to
kubernetes-csi-addons. This commit hence, removes
the volume replication sidecar and updates the
related documentation.
It will also update the csi-addons sidecar version
to the latest one.
Closes: #10655
Signed-off-by: yati1998 <ypadia@redhat.com>
The helm integration test was immediately removing both the cluster
and the operator charts, without waiting for any resource deletion
between the deletions. This causes intermittent failures in the
test since the cluster CR and mon secret and configmap to never
be deleted because the operator will be stopped before their
finalizers can be removed. Now we wait for their finalizers
to be removed before removing the operator chart.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook will support the most recent seven K8s releases,
so we update the min version to 1.19 for Rook v1.10.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Long ago the gateway type for s3 was removed from the object
store CRD, and k8s 1.25 is now complaining about the obsolete
setting still included in the test.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
K8s 1.25 is not capable of creating PSPs, so we need to pick
up v1.9.10 as the base for the upgrade test before upgrading
to master.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the pending release of 1.10, the upgrade tests
in master should be testing the upgrade from 1.9
to the master branch.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When a test has a slash in it's name the log collection can fail due to
the "directory" not existing, this makes sure the slashes are replaced
by underscores.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
"nolock" mountOptions has been added to prevent the following error in ci:
Output: mount.nfs: rpc.statd is not running but is required for remote locking.
mount.nfs: Either use '-o nolock' to keep locks local, or start statd.
Signed-off-by: Rakshith R <rar@redhat.com>
Block deletion of CephFilesystems when there are any raw Ceph
subvolumegroups present that have subvolumes in them. Empty
subvolumegroups will not block deletion.
One important subvolume group is "csi" which is the default location
where CSI subvolumes are kept. If this group is empty, it means that
there are no PVCs created based on the CephFilesystem in question. This
also holds true if there are external consumers of the filesystem in
external cluster mode.
Similarly, if there are any subvolumegroups (for example "_nogroup",
which includes subvolumes in the filesystem root) that contain
manually-created subvolumes, Rook will also see this and block deletion.
This comes into play currently with manually-created NFS exports.
A work-in progress aims to create a Ceph-CSI NFS export provisioner
which will likely create subvolumes in the "csi" group as well. This
implementation will catch this case also.
Rook still checks for CephFilesystemSubVolumeGroups explicitly in
addition to the check added here. This is to ensure that even empty
groups will block deletion if they are created via this CR type.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
updating node-driver-registrar to v2.5.1 release
, nfsplugin to 4.0.0 and csi-snapshotter to v6.0.1
Note:- with snapshotter v6.0.1 v1beta1 snapshot CRD
is not supported. If you want betav1 snapshot support
please switch to v5.x.x snapshotter.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Support is removed for k8s for various limitations such as
priority classes not working and csi driver feature
incompleteness and missing snapshots. Documentation is updated
with the new min version of k8s 1.17 and also the tests are
updated to run on the min version of 1.17.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
By default, we should set the priority class to one of the built-in
priority class names to ensure that pods critical to the storage
will be able to remain running when resources are low. Otherwise,
critical rook pods could be evicted and affect many other pods
that rely on the storage to continue functioning. The options have
been available in the CRs, but until now we have just not set the
defaults in the examples. Critical rook components are now set to
the priority class system-node-critical if they are generally pinned
to a node, and system-cluster-critical if they are critical to the storage.
Some pods such as the operator and crash collector do not have a
default priority class set in the examples since they don't affect
the data path.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
finally, admission controller will be enabled default
without any script/manual step. But it still requires cert-manager
to be installed which I believe is already installed in clusters.
**Note**
Code doesn't return error it just logs the error since
we don't want to stop reconciling if the admission controller fails.
We can work on this once the admission controller is stable.
Signed-off-by: subhamkrai <srai@redhat.com>
For supporting features like service account authentication for vault
KMS , a service account account need to attach with pod.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
The aws go lang sdk needs value for region, it is set differently in
various part of current code. With PR the value is always `us-east-1` so
that it will work RGW server without any issues.
This reverts commit 280c29f330.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>