Commit Graph
103 Commits
Author SHA1 Message Date
Joshua Hoblitt 57b7eeec80 object: add httpClient param to object.NewS3Agent()
To allow the caller to pass in their own transport when testing.

Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2024-12-12 10:20:41 -07:00
ee8bcad49d rgw: add support for keystone auth + swift/s3
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>

* extend the API object specs for swift and keystone integration

* adapt rgw to the new go-ceph version

  - The parameter lists of the API call have changes, as parameters
    ignored by the RGW Admin Ops API are no longer serialized, therefore
    the mock has to be adapted.

  - There is now validation for the user keys that are passed to the
    User get API, therefore things failed when we had empty keys in our
    User proxy object.

* expand the reconcile loop for the swift and keystone integration

* fix minor mistakes in design document

* add env var to pass extra args to minikube

  Minikube decides CPU cores and memory automatically based on the
  available resources on the machine which may be insufficient to
  run rook. This commit adds an environment variable to add arbitrary
  arguments to the minikube command, so both can be specified if
  desired.

* integration tests for swift and keystone

  The new integration of swift or s3 and keystone support by rook
  does not have any integration tests yet.

  This commit introduces integration tests for swift and keystone. The
  tests are done against a minimal keystone setup (keystone container
  image from Yaook-project (https://yaook.cloud), sqlite as database
  backend, cert-manager and trust-manager for test certificate setup).

  To prevent hardcoded credentials, passwords are generated
  by the tests. The integration tests use the openstack client
  (keystone- and swift-functionality) (https://docs.openstack.org/
  python-openstackclient/ latest/). This was a concious design decision
  to use client tooling as close as possible to the end user instead of
  using other go-libraries (such as gophercloud).

* add documentation on swift and keystone

  Currently there is no documentation on the use of Swift to access
  an object store as well as the use of OpenStack keystone for
  authentication.

  This commit adds documentation on the use of Swift and OpenStack
  keystone, as well as CRD-related documentation and an example setup.

* add integration tests for S3 via keystone

  This commit introduces integration tests for s3 and keystone. The
  tests are run against the same minimal keystone setup that the tests
  for swift and keystone use.

  The integration tests use the aws s3 client to use client tooling as
  close as possible to the end user instead of using other go-libraries.

Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
2024-08-08 14:26:21 +02:00
sp98 f6b1449faa core: cephblockpoolRadosNamespace cleanup
Clean up pool images and snapshots in the
radosnamespace

Signed-off-by: sp98 <sapillai@redhat.com>
2024-04-12 21:49:49 +05:30
parth-gr 60e879050a ci: delete svg in helm test
The filesystem is not being deleted because of the existing svg.
Add a cliet call to delete the deafult csi svg,
for helm test

Signed-off-by: parth-gr <partharora1010@gmail.com>
2023-12-12 00:52:19 +05:30
Anthony D'Atri 59d0240676 doc: improve ceph-csi-drivers.md and lintrolling
Signed-off-by: Anthony D'Atri <anthonyeleven@users.noreply.github.com>
2023-12-01 16:10:11 -07:00
parth-gr 46c241433d object: improve the error handling for multisite objs
here is the https://go.dev/play/p/SS9Q-dAiIx3 example which says the
error handling was wrongly implemented

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-14 15:12:15 +05:30
Redouane Kachach b3dd74ea20 docs: fixing some spelling issues
closes: https://github.com/rook/rook/issues/12987

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-10-03 13:50:17 +02:00
subhamkrai 0ed91458dc test: collect kube-system logs for debugging
added kube-system namespace also to collect
logs from to debug the smoke suite issue

Signed-off-by: subhamkrai <srai@redhat.com>
2023-07-26 14:09:15 +05:30
Jiffin Tony Thottan b48dc8a335 object: intial cosi driver controller design
Adding CephCOSIDriver CRD and controller. The controller will bring up
the ceph cosi driver when first object store is created in the rook
operator namespace. Then admin can defined COSI CRDs like BucketClass
and BucketAccessClass for different object stores deployed via Rook.
Using the BucketClass and BucketAccessClass, user can define
BucketAccess for backend bucket in the RGW. The CephCOSIDriver CRD
defines configuration options for ceph cosi driver. In the first version
its usability is minimal. Even if it is not defined Rook will bring up
the ceph cosi driver with default values.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-07-18 22:49:41 +05:30
Rakshith R db56ba22d2 ci: add nfs snap & restore e2e testcases
Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-18 15:06:58 +05:30
Rakshith R be10dac98c ci: enable and fixes for nfs ci
This commit anabled nfs csi ci and
add fixes/improvements to it like the
following:
- verify deletion of cephnfs and .nfs pool before proceeding
- verify pv deletion
- do not enable rook module
- reduce activeCount to 1 to save resources
- run cephnfs ci before cephfs ci since it cephfs
  ci is more resource intensive.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-10-06 11:53:46 +00:00
Liang Zheng 9a78f6af07 osd: update ceph status parse
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
2022-09-20 12:48:53 +08:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Rakshith R 5b017cfd67 ci: add tests for nfs csi pvc
This commit adds nfs csi pvc test into
ceph smoke suite.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-06-20 15:01:22 +05:30
Blaine Gardner a63844f8bf file: block deletion on more dependents
Block deletion of CephFilesystems when there are any raw Ceph
subvolumegroups present that have subvolumes in them. Empty
subvolumegroups will not block deletion.

One important subvolume group is "csi" which is the default location
where CSI subvolumes are kept. If this group is empty, it means that
there are no PVCs created based on the CephFilesystem in question. This
also holds true if there are external consumers of the filesystem in
external cluster mode.

Similarly, if there are any subvolumegroups (for example "_nogroup",
which includes subvolumes in the filesystem root) that contain
manually-created subvolumes, Rook will also see this and block deletion.
This comes into play currently with manually-created NFS exports.

A work-in progress aims to create a Ceph-CSI NFS export provisioner
which will likely create subvolumes in the "csi" group as well. This
implementation will catch this case also.

Rook still checks for CephFilesystemSubVolumeGroups explicitly in
addition to the check added here. This is to ensure that even empty
groups will block deletion if they are created via this CR type.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-06-14 13:02:35 -06:00
Jiffin Tony Thottan fc2b8012c6 object: use us-east-1 for aws go lang sdk
The aws go lang sdk needs value for region, it is set differently in
various part of current code. With PR the value is always `us-east-1` so
that it will work RGW server without any issues.

This reverts commit 280c29f330.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-03-24 14:37:39 +05:30
Jiffin Tony Thottan 438cf6abf1 test: update bucket notification integration test with http server
Add test cases to check notification is received by http server. For
this a sample http server https://github.com/thotz/pythonwebserver is
used.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-02-15 12:59:58 +05:30
Jiffin Tony Thottan a97747cece rgw: inject tls certs for bucket notification and topic operations
The certs for accessing TLS enabled RGW is saved as secrets and inject
them if controllers for notification and topics if request is sent to
TLS enabled RGW endpoint.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
2022-01-24 22:39:57 +05:30
Travis Nielsen 8fb758fee9 test: implement helm upgrade integration test
The helm tests were previously only for new installs, and did
not have an upgrade path. Now the upgrade path is tested
to give confidence in the helm upgrades.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-01-12 07:19:42 -07:00
Travis Nielsen 7ee9cc9d56 pool: allow configuration of built-in pools with non-k8s names
The built-in pools device_health_metrics and .nfs created by ceph
need to be configured for replicas, failure domain, etc.
To support this, we allow the pool to be created as a CR.
Since K8s does not support underscores in the resource names
the operator must translate this special pool name into
the name expected by ceph.

This also sets the basis for allowing filesystem data
pools to specify the desired pool name instead of requiring
a generated name.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-15 15:32:02 -07:00
Jiffin Tony Thottan 7a9a7123a0 test: add more test cases for bucket notfication
Added following test cases for bucket notification integration test
suite:

* different order: OBC - Topic - Notification
* different order: OBC - Notification - Topic
* adding a label to an existing OBC
* deleting a label from an existing OBC

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-12-02 23:02:45 +05:30
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Yuval Lifshitz 71ed45b69b rgw: implement bucket notifications for object storage
following the design from here:
https://github.com/rook/rook/blob/master/design/ceph/object/ceph-bucket-notification-crd.md

Closes: https://github.com/rook/rook/issues/5313
Signed-off-by: Yuval Lifshitz <ylifshit@redhat.com>
2021-11-04 11:20:40 +02:00
Jiffin Tony Thottan ca43800119 ceph: add options for cephobjectstore user
Adding options for quota, bucket limit, caps for the
`cephobjectstoreuser`.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-07 22:43:09 +05:30
Jiffin Tony Thottan f4bb47e440 ceph: add support for update() from lib-bucket-provisioner
Recently lib-bucket-provisioner add support for update() API.
Include that on the obc implementation since it can be used to
update quota for OBC.

Fixes: #7146

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-08-19 23:57:11 +05:30
Jiffin Tony Thottan 9e3cf68d04 test: ci test for TLS objectstore
Extend the object store smoke test to include TLS configurations.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-07-01 18:50:27 +05:30
Travis Nielsen c23238cddb ceph: refactor integration tests for simplification
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:10 -06:00
Blaine Gardner 9b0ba6ae8b ceph: add obc to upgrade test
Add object bucket claim to Ceph upgrade test.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-08 09:56:00 -07:00
Blaine Gardner 40fd80cf14 ceph: update smoke test to verify obc is bound
When verifying OBC creation, validating that secret and configmap exist
is good, but the definitive validation is to check that the OBC's phase
is "Bound".

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-07 13:58:22 -07:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai f9fafe62d4 ceph: handle golangci-lint linter unused
this commit will enable one more linter
in golangci-lint.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 22:53:24 +05:30
Madhu Rajanna e864b8d42d ceph: add E2E testing for snapshot and clone
Added E2E testing to create,delete and restore
a snapshot, create a pvc-pvc clone, install
and uninstall snapshot controller and snapshot
CRD.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-09-10 22:05:15 +05:30
Jiffin Tony Thottan f4240cb31d ceph: add quota support for obc
Closes: #5274
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-07-28 12:00:11 +05:30
Sébastien Han e517ac96aa ceph: fail if the pool fails to be deleted
Let's assert if the pool is still present.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-22 17:25:15 +02:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Travis Nielsen e1a9a9039b ceph: pool resize test hangs during cleanup
In Octopus 15.2.2 there is an rbd change that causes the rbd ls command
to hang if there are not sufficient OSDs to meet the replica requirement.
The smoke suite just started hanging today since the 15.2.2 image was
released with this change. Therefore, the test that resizes the pool
to replica 3 must resize back to 1 before cleaning up.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-19 22:02:37 -06:00
Satoru Takeuchi c0c0dc4ed9 tests: return bool if the functions are prefixed by "Is"
There are several functions that are prefixed by "Is" and return
just error. It's straightfoward to return bool to make the meaning
of these functions clearer.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-05-05 22:55:01 +09:00
Sébastien Han 2bd017102b ci: add retry when deleting rbd image
Sometimes the image still has watchers so let's retry the deletion.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-15 09:12:57 +02:00
Sébastien Han c1b8a7aa9e ceph: extract rbd-mirror to its own crd
Previously, the rbd-mirror daemon was integrated into the `CephCluster`
CRD. This wasn't really practical since we would have to wait for the
whole orchestration to be done to actually set it up. The same goes for
any CR update. Let's say you want to change the number of daemons, Rook
would go through mons, mgrs and osds until it get to rbd-mirror.
This triggers an undesired full orchestration.

With its own CRD this component just gains a lot more flexibility.

Closes: https://github.com/rook/rook/issues/5084
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-15 09:12:56 +02:00
Sébastien Han ca0a30f38d ceph: convert Filesystem controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-19 23:34:58 +01:00
Travis Nielsen 18b0e7d295 ceph: scrub ceph commands to write actions to the log
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Travis Nielsen ce1fedf884 ci: only list the images in the needed pool during cleanup
The pool cleanup only needs to happen for an individual pool.
No need to query the block images in all pools. One of the rgw
pools is periodically causing a hang when it is queried,
but there is no need to query for it when we are cleaning
up the pool tests.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-17 15:28:17 -06:00
Sébastien Han fd33d64b4f ci: do not hide pool deletion error
Print the error if any.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-06 14:35:23 +01:00
Travis Nielsen 65f3f84ab0 tests: ensure pools are purged during integration tests
With a finalizer on the pools, the pools were not always being purged
during the integration tests. Now the multicluster suite will ensure
its pool is purged.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-03 19:44:23 -07:00
Travis Nielsen a1db52c5a0 ceph: move integration test to csi driver
The integration tests have been mostly running on the flex driver
with only a newer test on the csi driver. With the CSI driver being
the preferred driver going forward, now the integration tests will
all be running with the CSI driver with the exception of a test
suite that is only dedicated to the flex driver.

A number of other test improvements are also made for code
readability, test stability, and removing unused options.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-17 16:24:28 -07:00
Travis Nielsen babd65d131 ceph: simplify most integration tests to single mds
The integration tests were always running two MDS daemons active,
with two standby. This is now parameterized so the test can request
how many MDS daemons to run. The smoke suite will run two active
and the rest will just run a single active MDS.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-04 12:17:27 -07:00
Mateusz Los cd713b11cb ceph: client CRD test fixes
wait for client to be updated before verifying caps

Signed-off-by: Mateusz Los <los.mateusz@gmail.com>
2019-12-16 23:06:33 +01:00
Mateusz Los 50f3d29cf4 ceph: client controller tests and docs
ceph client CRD refactoring
Add info about client crd to documentation
Add tests for client controller CRD

Signed-off-by: Mateusz Los <los.mateusz@gmail.com>
2019-12-09 12:48:44 +01:00
Travis Nielsen e32729269f ceph: clean up block integration tests more reliably
The block integration tests are a source of intermittent failures
in the CI. The PVCs were not being confirmed as removed before the
pools were deleted. Then the pool would not be removed since there
might still be a block image from the PVC that wasn't deleted yet.
Now the tests will ensure the PVCs and their block images are
removed before the pool is deleted.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2019-10-30 15:37:35 -06:00