Commit Graph
185 Commits
Author SHA1 Message Date
Steven Kreitzer 463d9fb420 csi: csi-snapshotter flag typo; upgrade csi-snapshotter
Signed-off-by: Steven Kreitzer <skre@skre.me>
2024-12-18 09:09:29 -06:00
df511fb58f ci: update golangci-lint to the latest version (v1.62)
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.

Additionally, it  silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
 and string format errors found by golangci-lint, while at it.

Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
2024-12-14 14:47:30 +01:00
Travis Nielsen b52ba6baca test: wait for mon daemons rather than mon canaries
The mon canaries may be created even when the mon daemons
are not created thereafter during the integration tests.
Therefore, the integration tests need to also query a label
specific to the mon daemon so the canaries are not a distraction
to the test.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-10-25 09:17:02 -06:00
Travis Nielsen 7ed77ddd13 core: remove obsolete creation of v1beta1 pruner cron jobs
The v1beta1 cron jobs have been obsolete since K8s 1.21,
and Rook has not supported that version of K8s
for many moons, so we can remove the obsolete code
for the handling of v1beta1 cron jobs for crash pruning.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-10-14 17:13:06 -06:00
Praveen M a1ddf4535d csi: update csi sidecars' image version
Below csi sidecars are updated with latest available versions

csi-resizer: v1.11.1
csi-provisioner: v5.0.1
csi-attacher: v4.6.1
csi-snapshotter: v8.0.1
csi-node-driver-registrar: v2.11.1

Signed-off-by: Praveen M <m.praveen@ibm.com>
2024-08-20 22:08:19 +05:30
ee8bcad49d rgw: add support for keystone auth + swift/s3
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>

* extend the API object specs for swift and keystone integration

* adapt rgw to the new go-ceph version

  - The parameter lists of the API call have changes, as parameters
    ignored by the RGW Admin Ops API are no longer serialized, therefore
    the mock has to be adapted.

  - There is now validation for the user keys that are passed to the
    User get API, therefore things failed when we had empty keys in our
    User proxy object.

* expand the reconcile loop for the swift and keystone integration

* fix minor mistakes in design document

* add env var to pass extra args to minikube

  Minikube decides CPU cores and memory automatically based on the
  available resources on the machine which may be insufficient to
  run rook. This commit adds an environment variable to add arbitrary
  arguments to the minikube command, so both can be specified if
  desired.

* integration tests for swift and keystone

  The new integration of swift or s3 and keystone support by rook
  does not have any integration tests yet.

  This commit introduces integration tests for swift and keystone. The
  tests are done against a minimal keystone setup (keystone container
  image from Yaook-project (https://yaook.cloud), sqlite as database
  backend, cert-manager and trust-manager for test certificate setup).

  To prevent hardcoded credentials, passwords are generated
  by the tests. The integration tests use the openstack client
  (keystone- and swift-functionality) (https://docs.openstack.org/
  python-openstackclient/ latest/). This was a concious design decision
  to use client tooling as close as possible to the end user instead of
  using other go-libraries (such as gophercloud).

* add documentation on swift and keystone

  Currently there is no documentation on the use of Swift to access
  an object store as well as the use of OpenStack keystone for
  authentication.

  This commit adds documentation on the use of Swift and OpenStack
  keystone, as well as CRD-related documentation and an example setup.

* add integration tests for S3 via keystone

  This commit introduces integration tests for s3 and keystone. The
  tests are run against the same minimal keystone setup that the tests
  for swift and keystone use.

  The integration tests use the aws s3 client to use client tooling as
  close as possible to the end user instead of using other go-libraries.

Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
2024-08-08 14:26:21 +02:00
Travis Nielsen 043b446675 tests: retry helm upgrade during upgrade tests
The helm upgrade tests have been failing frequently, but not
always, on the oldest version of K8s that is tested in the CI
for the past few months. Add a retry to attempt to get
the CI passing more consistently.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-07-24 08:14:21 -06:00
Praveen M faf7837621 csi: update csi sidecars' image version
Below sidecars are updated with latest available versions

csi-node-driver-registrar: v2.10.1
csi-resizer: v1.10.1
csi-provisioner: v4.0.1
csi-attacher: v4.5.1
csi-snapshotter: v7.0.2

Signed-off-by: Praveen M <m.praveen@ibm.com>
2024-04-25 18:58:00 +05:30
Madhu Rajanna 0d5bd70194 csi: install vgs CRD in tests
update the snapshot controller to 7.0.1
and install new Volumegroup CRD's

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2024-03-01 09:15:17 +01:00
Praveen M a80396df1b csi: update cmdline args as used by ceph-csi
This commit adds cmdline args to enable
1. RecoverVolumeExpansionFailure
2. PreventVolumeModeConversion
3. HonorPVReclaimPolicy

Signed-off-by: Praveen M <m.praveen@ibm.com>
2024-01-12 15:24:07 +05:30
parth-gr 40295a989c ci: fix objectsuite flakiness
objectstore deletion was failing with not found error,
Could Not get resource in k8s -- Failed to run:
kubectl [get -n object-ns CephObjectStore
other-tls-test-store -o json]
So added a check if it not found then donot check its
further condition

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-08 15:33:03 +05:30
Eng Zer Jun 77bff6a07c core: remove redundant len check
From the Go specification [1]:

  "1. For a nil slice, the number of iterations is 0."
  "3. If the map is nil, the number of iterations is 0."

`len` returns 0 if the slice or map is nil [2]. Therefore, checking
`len(v) > 0` before a loop is unnecessary.

[1]: https://go.dev/ref/spec#For_range
[2]: https://pkg.go.dev/builtin#len

Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
2023-10-05 20:16:46 +08:00
subhamkrai 1bd4ab9d81 ci: skip mgr pod restart count upgrade 1.22.x suite
for now, let's skip the mgr pod restart count
for upgrade suite 1.22.x to get the CI green
and so that we don't skip any other error in name
of mgr restart count.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-02 20:03:08 +05:30
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
parth-gr a84daf9bf0 core: change io/ioutil package to use io and os package
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues

Signed-off-by: parth-gr <paarora@redhat.com>
2023-02-17 20:38:29 +05:30
parth-gr 9319eaa86d ci: watch for pods that restart unexpectedly
will check if the podrestartcount is greater than 1,
If it is we will alert it and fail the CI
It is important to understand intermittent failures
to avoid too many false positives

Closes: https://github.com/rook/rook/issues/11380
Signed-off-by: parth-gr <paarora@redhat.com>
2022-12-14 21:50:15 +05:30
Travis Nielsen b693a5cca5 core: refactor crash collector for more node daemons
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md

Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-12-09 16:57:05 -07:00
Blaine Gardner fc56108865 object: allow status endpoints to be null
For the upgrade case, Rook can fail to set status information
on CephObjectStores after CRDs are updated but before the operator is
updated. To fix this, merely allow the slices to be null in the
CephObjectStore's status.endpoints.

This PR seems to be aggravating the helm filesystem upgrade test. Allow
30 more seconds for CRDs to be done before bailing. In debugging, the
filesystem often needed only an extra 3 seconds to succeed.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-08 13:36:15 -07:00
Rakshith R be10dac98c ci: enable and fixes for nfs ci
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
  ci is more resource intensive.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-06 11:53:46 +00:00
Liang Zheng f9c360e609 osd: optimize device probe
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.

Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2022-09-20 10:41:51 +08:00
Alexander Trost 49440e8b20 test: fix test pod log collector
When a test has a slash in it's name the log collection can fail due to
the "directory" not existing, this makes sure the slashes are replaced
by underscores.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-08-12 23:44:13 +02:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Rakshith R 034c4756d3 ci: make sure PVC is deleted before proceeding for file test
Signed-off-by: Rakshith R <rar@redhat.com>
2022-07-04 11:36:37 +05:30
Blaine Gardner a63844f8bf file: block deletion on more dependents
Block deletion of CephFilesystems when there are any raw Ceph
subvolumegroups present that have subvolumes in them. Empty
subvolumegroups will not block deletion.

One important subvolume group is "csi" which is the default location
where CSI subvolumes are kept. If this group is empty, it means that
there are no PVCs created based on the CephFilesystem in question. This
also holds true if there are external consumers of the filesystem in
external cluster mode.

Similarly, if there are any subvolumegroups (for example "_nogroup",
which includes subvolumes in the filesystem root) that contain
manually-created subvolumes, Rook will also see this and block deletion.
This comes into play currently with manually-created NFS exports.

A work-in progress aims to create a Ceph-CSI NFS export provisioner
which will likely create subvolumes in the "csi" group as well. This
implementation will catch this case also.

Rook still checks for CephFilesystemSubVolumeGroups explicitly in
addition to the check added here. This is to ensure that even empty
groups will block deletion if they are created via this CR type.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-06-14 13:02:35 -06:00
Travis Nielsen 2db3910052 test: remove dead test helper code
The K8sHelper has a number of methods that are no longer
in use that we can remove.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-25 15:28:29 -06:00
Travis Nielsen fb86955f01 core: examples set default priority class names
By default, we should set the priority class to one of the built-in
priority class names to ensure that pods critical to the storage
will be able to remain running when resources are low. Otherwise,
critical rook pods could be evicted and affect many other pods
that rely on the storage to continue functioning. The options have
been available in the CRs, but until now we have just not set the
defaults in the examples. Critical rook components are now set to
the priority class system-node-critical if they are generally pinned
to a node, and system-cluster-critical if they are critical to the storage.
Some pods such as the operator and crash collector do not have a
default priority class set in the examples since they don't affect
the data path.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-19 16:09:34 -06:00
Divyansh Kamboj 9008409f87 core: add context parameter to functions
This commit adds context parameter to various functions, and remove the
usage of context.TODO.

Closes: https://github.com/rook/rook/issues/8701
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2022-03-22 08:19:07 +05:30
Madhu Rajanna 97e5b78c11 csi: cover error cases in snapshot ready check
Currently, we sleep only if the snapshot is not ready
and try again, this covers only the happy path.
the checks might fail to get the snapshot object
or the `.status.readyToUse` might not be set yet.
First sleep and then try to check snapshot is
ready or not, if we do this we cover all the cases
when checking the snapshot ready status.
Changed sleep internal to do incremental sleep to
give more time for tests.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-01-31 13:07:37 +05:30
Madhu Rajanna c40ff24410 csi: skip validation when installing snapshotter
in kubernetes 1.18 we are seeing CRD validation
errors when installing with kubectl, adding
validate=false to skip the validation and
to make tests works with kubernetes 1.18

closes: #9670

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-01-28 08:25:07 +05:30
Madhu Rajanna 53e12d687f csi: check deployment for snapshot controller
The external-snapshotter was deployed as statefulset
in 4.x and now its deployed as a deployment. updated
the check in CI to make sure deployment is created.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-01-27 23:06:16 +05:30
Madhu RajannaandMathieu Parent cf46615688 csi: bump csi snapshotter image to v5
updating the csi-snapshotter and dependencies
to v5.0.1 released version.

Co-authored-by: Mathieu Parent <mathieu.parent@insee.fr>
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-01-27 21:17:51 +05:30
Travis Nielsen 741f211a13 helm: apply operator settings in configmap instead of deployment
Some operator settings can be applied dynamically in a configmap
instead of requiring the operator to restart. While these settings
may be infrequently updated, applying these settings in the configmap
can avoid an unnecessary operator restart. This will also make the
operator deployment more consistent with the non-helm install.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-01-20 14:41:28 -07:00
Travis Nielsen 8fb758fee9 test: implement helm upgrade integration test
The helm tests were previously only for new installs, and did
not have an upgrade path. Now the upgrade path is tested
to give confidence in the helm upgrades.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-01-12 07:19:42 -07:00
Jiffin Tony Thottan 7a9a7123a0 test: add more test cases for bucket notfication
Added following test cases for bucket notification integration test
suite:

* different order: OBC - Topic - Notification
* different order: OBC - Notification - Topic
* adding a label to an existing OBC
* deleting a label from an existing OBC

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-12-02 23:02:45 +05:30
Sébastien Han c890710b63 core: change directory layout
As per discussion, proposing a new layout for the charts/yaml/olm files.

./deploy
├── charts
│   ├── rook-ceph
│   │   └── templates
│   └── rook-ceph-cluster
│       └── templates
├── examples
│   ├── csi
│   │   ├── cephfs
│   │   └── rbd
│   ├── flex
│   ├── monitoring
│   ├── pre-k8s-1.16
└── olm
    └── assemble

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-30 09:12:53 +01:00
Sébastien Han b4a36e9967 ci: wait longer for pod label to be deleted
I've seen cases were the CI needs a few more seconds to delete and
object store. When logging in the runner, the object store is gone and
the timing matches too with the runner's logs (comparing with the
operator's logs).

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-28 15:24:39 +02:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Travis Nielsen 07ddcda0ce ceph: update test to watch for v1 cronjob
The v1beta1 cronjob is deprecated and needs to use v1 for K8s 1.16 and newer

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-22 11:25:55 -06:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Andy Bursavich 11f0add706 build: update golangci-lint version and nolint comment format
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:50:04 -07:00
Andy Bursavich 7ab1baa794 test: extract utils dependency on k8s.io/kubernetes
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:36:08 -07:00
subhamkrai 0ff452fdd1 ceph: remove unnecessary file
this commit removes `mysql_helper.go`,
unnecessary file.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-06-15 15:28:13 +05:30
Henry Zhang ecf6bdcfae ceph: add build and tests for rook-ceph-cluster Helm chart
Adds rook-ceph-cluster chart to chart build, add tests.
Pulled out some common functionality between the Helm and non-Helm installers

Signed-off-by: Henry Zhang <me@henry.dev>
2021-05-27 01:08:27 -07:00
Rakshith R 7987e1b8a2 ceph: update snapshot APIs from v1beta1 to v1
This commit updates external-snapshotter version to
v4.0.0 which supports snapshots v1.
Rook now defaults to enabling RBD and CephFS snapshotter
for K8s >= v1.17 and disabling it for K8s <= v1.16.
Supporting changes in documents and examples yaml files
are made.

Signed-off-by: Rakshith R <rar@redhat.com>
2021-05-05 21:04:10 +05:30
Yannis Zarkadas b71f1077b5 test: replace in-house config loader for controller-runtime's
Rook's testing utilities include a function for loading the default
kubeconfig for a cluster. This function requires maintainance effort and
also doesn't support authentication methods like Basic Authentication.
Instead of adding support for it, drop the config loader and use the one
provided by the controller-runtime library.

Signed-off-by: Yannis Zarkadas <yanniszark@arrikto.com>
2021-04-14 21:40:50 +03:00
Travis Nielsen bdbf264e68 ceph: integration tests in github actions skip cleanup
The integration tests run independently in the github actions so there is
no need to cleanup from every test. The cleanup is still needed in the
Jenkins tests where all the suites run serially.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:11 -06:00
Travis Nielsen c23238cddb ceph: refactor integration tests for simplification
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:10 -06:00
Blaine Gardner 9b0ba6ae8b ceph: add obc to upgrade test
Add object bucket claim to Ceph upgrade test.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-08 09:56:00 -07:00