Commit Graph
175 Commits
Author SHA1 Message Date
parth-gr 40295a989c ci: fix objectsuite flakiness
objectstore deletion was failing with not found error,
Could Not get resource in k8s -- Failed to run:
kubectl [get -n object-ns CephObjectStore
other-tls-test-store -o json]
So added a check if it not found then donot check its
further condition

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-08 15:33:03 +05:30
Eng Zer Jun 77bff6a07c core: remove redundant len check
From the Go specification [1]:

  "1. For a nil slice, the number of iterations is 0."
  "3. If the map is nil, the number of iterations is 0."

`len` returns 0 if the slice or map is nil [2]. Therefore, checking
`len(v) > 0` before a loop is unnecessary.

[1]: https://go.dev/ref/spec#For_range
[2]: https://pkg.go.dev/builtin#len

Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
2023-10-05 20:16:46 +08:00
subhamkrai 1bd4ab9d81 ci: skip mgr pod restart count upgrade 1.22.x suite
for now, let's skip the mgr pod restart count
for upgrade suite 1.22.x to get the CI green
and so that we don't skip any other error in name
of mgr restart count.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-02 20:03:08 +05:30
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
parth-gr a84daf9bf0 core: change io/ioutil package to use io and os package
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues

Signed-off-by: parth-gr <paarora@redhat.com>
2023-02-17 20:38:29 +05:30
parth-gr 9319eaa86d ci: watch for pods that restart unexpectedly
will check if the podrestartcount is greater than 1,
If it is we will alert it and fail the CI
It is important to understand intermittent failures
to avoid too many false positives

Closes: https://github.com/rook/rook/issues/11380
Signed-off-by: parth-gr <paarora@redhat.com>
2022-12-14 21:50:15 +05:30
Travis Nielsen b693a5cca5 core: refactor crash collector for more node daemons
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md

Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-12-09 16:57:05 -07:00
Blaine Gardner fc56108865 object: allow status endpoints to be null
For the upgrade case, Rook can fail to set status information
on CephObjectStores after CRDs are updated but before the operator is
updated. To fix this, merely allow the slices to be null in the
CephObjectStore's status.endpoints.

This PR seems to be aggravating the helm filesystem upgrade test. Allow
30 more seconds for CRDs to be done before bailing. In debugging, the
filesystem often needed only an extra 3 seconds to succeed.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-11-08 13:36:15 -07:00
Rakshith R be10dac98c ci: enable and fixes for nfs ci
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
  ci is more resource intensive.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-06 11:53:46 +00:00
Liang Zheng f9c360e609 osd: optimize device probe
1. Eliminate possible memory leaks of timer.
2. Eliminate duplicated events between udev events and kernel events.
3. Empty struct have the lowest size.

Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2022-09-20 10:41:51 +08:00
Alexander Trost 49440e8b20 test: fix test pod log collector
When a test has a slash in it's name the log collection can fail due to
the "directory" not existing, this makes sure the slashes are replaced
by underscores.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-08-12 23:44:13 +02:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Rakshith R 034c4756d3 ci: make sure PVC is deleted before proceeding for file test
Signed-off-by: Rakshith R <rar@redhat.com>
2022-07-04 11:36:37 +05:30
Blaine Gardner a63844f8bf file: block deletion on more dependents
Block deletion of CephFilesystems when there are any raw Ceph
subvolumegroups present that have subvolumes in them. Empty
subvolumegroups will not block deletion.

One important subvolume group is "csi" which is the default location
where CSI subvolumes are kept. If this group is empty, it means that
there are no PVCs created based on the CephFilesystem in question. This
also holds true if there are external consumers of the filesystem in
external cluster mode.

Similarly, if there are any subvolumegroups (for example "_nogroup",
which includes subvolumes in the filesystem root) that contain
manually-created subvolumes, Rook will also see this and block deletion.
This comes into play currently with manually-created NFS exports.

A work-in progress aims to create a Ceph-CSI NFS export provisioner
which will likely create subvolumes in the "csi" group as well. This
implementation will catch this case also.

Rook still checks for CephFilesystemSubVolumeGroups explicitly in
addition to the check added here. This is to ensure that even empty
groups will block deletion if they are created via this CR type.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-06-14 13:02:35 -06:00
Travis Nielsen 2db3910052 test: remove dead test helper code
The K8sHelper has a number of methods that are no longer
in use that we can remove.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-25 15:28:29 -06:00
Travis Nielsen fb86955f01 core: examples set default priority class names
By default, we should set the priority class to one of the built-in
priority class names to ensure that pods critical to the storage
will be able to remain running when resources are low. Otherwise,
critical rook pods could be evicted and affect many other pods
that rely on the storage to continue functioning. The options have
been available in the CRs, but until now we have just not set the
defaults in the examples. Critical rook components are now set to
the priority class system-node-critical if they are generally pinned
to a node, and system-cluster-critical if they are critical to the storage.
Some pods such as the operator and crash collector do not have a
default priority class set in the examples since they don't affect
the data path.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-19 16:09:34 -06:00
Divyansh Kamboj 9008409f87 core: add context parameter to functions
This commit adds context parameter to various functions, and remove the
usage of context.TODO.

Closes: https://github.com/rook/rook/issues/8701
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2022-03-22 08:19:07 +05:30
Madhu Rajanna 97e5b78c11 csi: cover error cases in snapshot ready check
Currently, we sleep only if the snapshot is not ready
and try again, this covers only the happy path.
the checks might fail to get the snapshot object
or the `.status.readyToUse` might not be set yet.
First sleep and then try to check snapshot is
ready or not, if we do this we cover all the cases
when checking the snapshot ready status.
Changed sleep internal to do incremental sleep to
give more time for tests.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-01-31 13:07:37 +05:30
Madhu Rajanna c40ff24410 csi: skip validation when installing snapshotter
in kubernetes 1.18 we are seeing CRD validation
errors when installing with kubectl, adding
validate=false to skip the validation and
to make tests works with kubernetes 1.18

closes: #9670

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-01-28 08:25:07 +05:30
Madhu Rajanna 53e12d687f csi: check deployment for snapshot controller
The external-snapshotter was deployed as statefulset
in 4.x and now its deployed as a deployment. updated
the check in CI to make sure deployment is created.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-01-27 23:06:16 +05:30
Madhu RajannaandMathieu Parent cf46615688 csi: bump csi snapshotter image to v5
updating the csi-snapshotter and dependencies
to v5.0.1 released version.

Co-authored-by: Mathieu Parent <mathieu.parent@insee.fr>
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-01-27 21:17:51 +05:30
Travis Nielsen 741f211a13 helm: apply operator settings in configmap instead of deployment
Some operator settings can be applied dynamically in a configmap
instead of requiring the operator to restart. While these settings
may be infrequently updated, applying these settings in the configmap
can avoid an unnecessary operator restart. This will also make the
operator deployment more consistent with the non-helm install.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-01-20 14:41:28 -07:00
Travis Nielsen 8fb758fee9 test: implement helm upgrade integration test
The helm tests were previously only for new installs, and did
not have an upgrade path. Now the upgrade path is tested
to give confidence in the helm upgrades.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-01-12 07:19:42 -07:00
Jiffin Tony Thottan 7a9a7123a0 test: add more test cases for bucket notfication
Added following test cases for bucket notification integration test
suite:

* different order: OBC - Topic - Notification
* different order: OBC - Notification - Topic
* adding a label to an existing OBC
* deleting a label from an existing OBC

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-12-02 23:02:45 +05:30
Sébastien Han c890710b63 core: change directory layout
As per discussion, proposing a new layout for the charts/yaml/olm files.

./deploy
├── charts
│   ├── rook-ceph
│   │   └── templates
│   └── rook-ceph-cluster
│       └── templates
├── examples
│   ├── csi
│   │   ├── cephfs
│   │   └── rbd
│   ├── flex
│   ├── monitoring
│   ├── pre-k8s-1.16
└── olm
    └── assemble

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-30 09:12:53 +01:00
Sébastien Han b4a36e9967 ci: wait longer for pod label to be deleted
I've seen cases were the CI needs a few more seconds to delete and
object store. When logging in the runner, the object store is gone and
the timing matches too with the runner's logs (comparing with the
operator's logs).

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-28 15:24:39 +02:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Travis Nielsen 07ddcda0ce ceph: update test to watch for v1 cronjob
The v1beta1 cronjob is deprecated and needs to use v1 for K8s 1.16 and newer

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-22 11:25:55 -06:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Andy Bursavich 11f0add706 build: update golangci-lint version and nolint comment format
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:50:04 -07:00
Andy Bursavich 7ab1baa794 test: extract utils dependency on k8s.io/kubernetes
Signed-off-by: Andy Bursavich <abursavich@gmail.com>
2021-06-21 08:36:08 -07:00
subhamkrai 0ff452fdd1 ceph: remove unnecessary file
this commit removes `mysql_helper.go`,
unnecessary file.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-06-15 15:28:13 +05:30
Henry Zhang ecf6bdcfae ceph: add build and tests for rook-ceph-cluster Helm chart
Adds rook-ceph-cluster chart to chart build, add tests.
Pulled out some common functionality between the Helm and non-Helm installers

Signed-off-by: Henry Zhang <me@henry.dev>
2021-05-27 01:08:27 -07:00
Rakshith R 7987e1b8a2 ceph: update snapshot APIs from v1beta1 to v1
This commit updates external-snapshotter version to
v4.0.0 which supports snapshots v1.
Rook now defaults to enabling RBD and CephFS snapshotter
for K8s >= v1.17 and disabling it for K8s <= v1.16.
Supporting changes in documents and examples yaml files
are made.

Signed-off-by: Rakshith R <rar@redhat.com>
2021-05-05 21:04:10 +05:30
Yannis Zarkadas b71f1077b5 test: replace in-house config loader for controller-runtime's
Rook's testing utilities include a function for loading the default
kubeconfig for a cluster. This function requires maintainance effort and
also doesn't support authentication methods like Basic Authentication.
Instead of adding support for it, drop the config loader and use the one
provided by the controller-runtime library.

Signed-off-by: Yannis Zarkadas <yanniszark@arrikto.com>
2021-04-14 21:40:50 +03:00
Travis Nielsen bdbf264e68 ceph: integration tests in github actions skip cleanup
The integration tests run independently in the github actions so there is
no need to cleanup from every test. The cleanup is still needed in the
Jenkins tests where all the suites run serially.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:11 -06:00
Travis Nielsen c23238cddb ceph: refactor integration tests for simplification
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:10 -06:00
Blaine Gardner 9b0ba6ae8b ceph: add obc to upgrade test
Add object bucket claim to Ceph upgrade test.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-08 09:56:00 -07:00
Blaine Gardner 40fd80cf14 ceph: update smoke test to verify obc is bound
When verifying OBC creation, validating that secret and configmap exist
is good, but the definitive validation is to check that the OBC's phase
is "Bound".

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-07 13:58:22 -07:00
Travis Nielsen c15499ed7e ceph: use the correct snapshot controller version in the tests
The snapshot controller is specifying to use the canary image instead
of the expected version. Update the tests to set the correct version.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-16 10:25:00 -07:00
Renan Campos ad9dcb0c05 ceph: periodically prune crash entries older than user-provided days
Rook's crashcollector pod posts entries to the ceph cluster when a crash occurs.
Over time the number cluster may hold crash entries needlessly.
To clean up old crash entries, this PR adds a field to the ceph cluster CR for the user to specify the number of days a crash entry should be kept for.
Providing a value for the field keepXDays creates a cronjob that runs every day at midnight, calling "ceph crash prune <keepXDays>".

Closes: https://github.com/rook/rook/issues/6332
Signed-off-by: Renan Campos <rcampos@redhat.com>
2020-11-24 17:06:18 -05:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
subhamkrai 49a3511532 ci: manifests changes to run gh action
to run the integration test there needs
to be a few changes in manifests like using
`deviceFilter` and other related changes.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-11-18 09:31:07 +05:30
subhamkrai a2d3d1d8a4 ceph: update snapshotterVersion to v3.0.0
in the integration test we still use
snapshotteVersion v2.1.0 but it's expected to
use v3.0.0.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-11-16 12:10:50 +05:30
Pete Birley 152a05c85e ceph: update to helm 3 for the rook chart
This updates the chart to make use of helm3 which has been released
for some time, and also permits CRDs to be installed pror to other objects
allowing the chart to be deployed at the same time as CRs for rook objects.

Co-authored-by: Pete Birley <pete@port.direct>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-11 14:41:24 -07:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
subhamkrai cb0ca66a6b ci: enable gosec linter in golangci-lint
golangci-lint linter gosec showing more errors
than gosec gh. This commit resolve
new errors. And, removing
nosec comments from autogenerated files.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-10-19 11:50:06 +05:30
subhamkrai 0fddfcf307 ceph: handle golangci-lint linter errcheck error
this commit handle golangci-lint linter errcheck.

`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases

To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-30 22:24:34 +05:30