The ci was using a pretty old version og golangci-lint.
This updates to the latest version.
Additionally, it silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
and string format errors found by golangci-lint, while at it.
Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
The mon canaries may be created even when the mon daemons
are not created thereafter during the integration tests.
Therefore, the integration tests need to also query a label
specific to the mon daemon so the canaries are not a distraction
to the test.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The v1beta1 cron jobs have been obsolete since K8s 1.21,
and Rook has not supported that version of K8s
for many moons, so we can remove the obsolete code
for the handling of v1beta1 cron jobs for crash pruning.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>
* extend the API object specs for swift and keystone integration
* adapt rgw to the new go-ceph version
- The parameter lists of the API call have changes, as parameters
ignored by the RGW Admin Ops API are no longer serialized, therefore
the mock has to be adapted.
- There is now validation for the user keys that are passed to the
User get API, therefore things failed when we had empty keys in our
User proxy object.
* expand the reconcile loop for the swift and keystone integration
* fix minor mistakes in design document
* add env var to pass extra args to minikube
Minikube decides CPU cores and memory automatically based on the
available resources on the machine which may be insufficient to
run rook. This commit adds an environment variable to add arbitrary
arguments to the minikube command, so both can be specified if
desired.
* integration tests for swift and keystone
The new integration of swift or s3 and keystone support by rook
does not have any integration tests yet.
This commit introduces integration tests for swift and keystone. The
tests are done against a minimal keystone setup (keystone container
image from Yaook-project (https://yaook.cloud), sqlite as database
backend, cert-manager and trust-manager for test certificate setup).
To prevent hardcoded credentials, passwords are generated
by the tests. The integration tests use the openstack client
(keystone- and swift-functionality) (https://docs.openstack.org/
python-openstackclient/ latest/). This was a concious design decision
to use client tooling as close as possible to the end user instead of
using other go-libraries (such as gophercloud).
* add documentation on swift and keystone
Currently there is no documentation on the use of Swift to access
an object store as well as the use of OpenStack keystone for
authentication.
This commit adds documentation on the use of Swift and OpenStack
keystone, as well as CRD-related documentation and an example setup.
* add integration tests for S3 via keystone
This commit introduces integration tests for s3 and keystone. The
tests are run against the same minimal keystone setup that the tests
for swift and keystone use.
The integration tests use the aws s3 client to use client tooling as
close as possible to the end user instead of using other go-libraries.
Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
The helm upgrade tests have been failing frequently, but not
always, on the oldest version of K8s that is tested in the CI
for the past few months. Add a retry to attempt to get
the CI passing more consistently.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
objectstore deletion was failing with not found error,
Could Not get resource in k8s -- Failed to run:
kubectl [get -n object-ns CephObjectStore
other-tls-test-store -o json]
So added a check if it not found then donot check its
further condition
Signed-off-by: parth-gr <paarora@redhat.com>
From the Go specification [1]:
"1. For a nil slice, the number of iterations is 0."
"3. If the map is nil, the number of iterations is 0."
`len` returns 0 if the slice or map is nil [2]. Therefore, checking
`len(v) > 0` before a loop is unnecessary.
[1]: https://go.dev/ref/spec#For_range
[2]: https://pkg.go.dev/builtin#len
Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
for now, let's skip the mgr pod restart count
for upgrade suite 1.22.x to get the CI green
and so that we don't skip any other error in name
of mgr restart count.
Signed-off-by: subhamkrai <srai@redhat.com>
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
will check if the podrestartcount is greater than 1,
If it is we will alert it and fail the CI
It is important to understand intermittent failures
to avoid too many false positives
Closes: https://github.com/rook/rook/issues/11380
Signed-off-by: parth-gr <paarora@redhat.com>
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md
Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For the upgrade case, Rook can fail to set status information
on CephObjectStores after CRDs are updated but before the operator is
updated. To fix this, merely allow the slices to be null in the
CephObjectStore's status.endpoints.
This PR seems to be aggravating the helm filesystem upgrade test. Allow
30 more seconds for CRDs to be done before bailing. In debugging, the
filesystem often needed only an extra 3 seconds to succeed.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
ci is more resource intensive.
Signed-off-by: Rakshith R <rar@redhat.com>
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
When a test has a slash in it's name the log collection can fail due to
the "directory" not existing, this makes sure the slashes are replaced
by underscores.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
Block deletion of CephFilesystems when there are any raw Ceph
subvolumegroups present that have subvolumes in them. Empty
subvolumegroups will not block deletion.
One important subvolume group is "csi" which is the default location
where CSI subvolumes are kept. If this group is empty, it means that
there are no PVCs created based on the CephFilesystem in question. This
also holds true if there are external consumers of the filesystem in
external cluster mode.
Similarly, if there are any subvolumegroups (for example "_nogroup",
which includes subvolumes in the filesystem root) that contain
manually-created subvolumes, Rook will also see this and block deletion.
This comes into play currently with manually-created NFS exports.
A work-in progress aims to create a Ceph-CSI NFS export provisioner
which will likely create subvolumes in the "csi" group as well. This
implementation will catch this case also.
Rook still checks for CephFilesystemSubVolumeGroups explicitly in
addition to the check added here. This is to ensure that even empty
groups will block deletion if they are created via this CR type.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
By default, we should set the priority class to one of the built-in
priority class names to ensure that pods critical to the storage
will be able to remain running when resources are low. Otherwise,
critical rook pods could be evicted and affect many other pods
that rely on the storage to continue functioning. The options have
been available in the CRs, but until now we have just not set the
defaults in the examples. Critical rook components are now set to
the priority class system-node-critical if they are generally pinned
to a node, and system-cluster-critical if they are critical to the storage.
Some pods such as the operator and crash collector do not have a
default priority class set in the examples since they don't affect
the data path.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Currently, we sleep only if the snapshot is not ready
and try again, this covers only the happy path.
the checks might fail to get the snapshot object
or the `.status.readyToUse` might not be set yet.
First sleep and then try to check snapshot is
ready or not, if we do this we cover all the cases
when checking the snapshot ready status.
Changed sleep internal to do incremental sleep to
give more time for tests.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
in kubernetes 1.18 we are seeing CRD validation
errors when installing with kubectl, adding
validate=false to skip the validation and
to make tests works with kubernetes 1.18
closes: #9670
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
The external-snapshotter was deployed as statefulset
in 4.x and now its deployed as a deployment. updated
the check in CI to make sure deployment is created.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
updating the csi-snapshotter and dependencies
to v5.0.1 released version.
Co-authored-by: Mathieu Parent <mathieu.parent@insee.fr>
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Some operator settings can be applied dynamically in a configmap
instead of requiring the operator to restart. While these settings
may be infrequently updated, applying these settings in the configmap
can avoid an unnecessary operator restart. This will also make the
operator deployment more consistent with the non-helm install.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The helm tests were previously only for new installs, and did
not have an upgrade path. Now the upgrade path is tested
to give confidence in the helm upgrades.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Added following test cases for bucket notification integration test
suite:
* different order: OBC - Topic - Notification
* different order: OBC - Notification - Topic
* adding a label to an existing OBC
* deleting a label from an existing OBC
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
I've seen cases were the CI needs a few more seconds to delete and
object store. When logging in the runner, the object store is gone and
the timing matches too with the runner's logs (comparing with the
operator's logs).
Signed-off-by: Sébastien Han <seb@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.
So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.
Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.
It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.
Signed-off-by: Sébastien Han <seb@redhat.com>
Adds rook-ceph-cluster chart to chart build, add tests.
Pulled out some common functionality between the Helm and non-Helm installers
Signed-off-by: Henry Zhang <me@henry.dev>
This commit updates external-snapshotter version to
v4.0.0 which supports snapshots v1.
Rook now defaults to enabling RBD and CephFS snapshotter
for K8s >= v1.17 and disabling it for K8s <= v1.16.
Supporting changes in documents and examples yaml files
are made.
Signed-off-by: Rakshith R <rar@redhat.com>
Rook's testing utilities include a function for loading the default
kubeconfig for a cluster. This function requires maintainance effort and
also doesn't support authentication methods like Basic Authentication.
Instead of adding support for it, drop the config loader and use the one
provided by the controller-runtime library.
Signed-off-by: Yannis Zarkadas <yanniszark@arrikto.com>
The integration tests run independently in the github actions so there is
no need to cleanup from every test. The cleanup is still needed in the
Jenkins tests where all the suites run serially.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>