Commit Graph
362 Commits
Author SHA1 Message Date
Travis Nielsen 9d2aa1f6bd test: generate long node name depending on test suite
The generation of a long node name in the integration tests was
being done based on the k8s version. In the past, older K8s versions
did not support the changing name. Now it's more maintainable if
we generate the long name depending on the test suite.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-06 08:03:56 -07:00
Travis Nielsen d9ac8ce490 test: upgrade integration test from 1.7 to master
The upgrade integration test was from rook v1.6 to the latest master.
This was necessary until we are ready for the v1.8 release, from which
time we want to focus the upgrade testing from v1.7 to the latest
master.

The duplication in the test CRs and other resources is now reduced
by the upgrade calling a thin wrapper to forward a call to the
master version of the resource. When a new feature is added that
needs to be differentiated from the previous version, the method
then can be implemented instead of wrapping the master implementation.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-30 08:07:42 -07:00
Travis Nielsen ba54567e8f Merge pull request #9176 from TomHellier/9174-ingress-support-more-k8s-versions
helm: Allow further configurability of the ingress version
2021-11-23 13:47:33 -07:00
Tom Hellier ba44602477 helm: allow further configurability of ingress version
The ingress api version changed when it went to v1, and this has caused some upheaval
throughout the kubernetes ecosystem. This commit uses a common method of deciding which
ingress api to use, and allows the optional override of the kubernetes version
presented to helm using the helm build-in capabilities.
also add an ingress into the helm integration tests so any regressions to how ingresses
are handled in the future are caught easier.

Closes rook#9174

Signed-off-by: Tom Hellier <me@tomhellier.com>
2021-11-22 10:03:02 +00:00
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Yuichiro Ueno 4cc716a7ca core: add context parameter to k8sutil node
This commit adds context parameter to k8sutil node functions. By this,
we can handle cancellation during API call of node resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:45:59 +09:00
Sébastien Han ecd7fa7880 Merge pull request #9164 from y1r/add-context-k8sutil-pod
core: add context parameter to k8sutil pod
2021-11-15 11:58:36 +01:00
Yuichiro Ueno 0559977b8a core: add context parameter to k8sutil pod
This commit adds context parameter to k8sutil pod functions. By this, we
can handle cancellation during API call of pod resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 15:39:41 +09:00
Yuichiro Ueno 0b575703c7 core: add context parameter to k8sutil deployment
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 14:58:13 +09:00
Yuval Lifshitz 71ed45b69b rgw: implement bucket notifications for object storage
following the design from here:
https://github.com/rook/rook/blob/master/design/ceph/object/ceph-bucket-notification-crd.md

Closes: https://github.com/rook/rook/issues/5313
Signed-off-by: Yuval Lifshitz <ylifshit@redhat.com>
2021-11-04 11:20:40 +02:00
Yuzuki Mimura 536b59ef0f rgw: change the way to livenessProbe and introduce readinessProbe
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.

Closes: #8407

Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-10-29 15:07:29 +00:00
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
Travis Nielsen c420f2309c csi: no longer install the volumereplication crds from rook
The volume replication CRDs are an external component, not owned by Rook.
Therefore, they should be installed as any other independent component
in case the admin will install other consumers of the volumereplication CRDs
in the future in addition to Rook and the CSI driver.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-06 16:27:44 +00:00
Blaine Gardner 08cb6785a4 Merge pull request #8804 from jmolmo/fix_timeout
ceph: retry <ceph orch> commands when they fail
2021-10-01 10:55:32 -06:00
Sébastien Han cda5dad291 rgw: use insecure TLS for bucket health check
We have seen cases where the signed certificate used for the RGW does not
contain the internal DNS endpoint, resulting in the health check to fail
since the certificate is not valid for this domain.
People consuming the gateways by external clients and for specific
domains do not necessarily have the internal DNS configured in the
certificate.
So let's be a bit more flexible and simply ensure a connectivity check
and bypass the certificate validation.

Also, this is fixing the tls code in newS3Agent and adds unit tests.

Closes: #8663
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-28 14:48:07 +02:00
Sébastien Han d2e221c206 Merge pull request #8850 from leseb/fix-mgr-host
ci: fix mgr test for host
2021-09-28 10:18:01 +02:00
Sébastien Han 5d8a1deda1 ci: fix mgr test for host
The `mgr host ls` looks like:

`"addr": "10.1.0.27/fv-az244-362", "hostname": "fv-az244-362"`

Previously, we were picking `addr` which contains both IP and hostname
where the k8snode has the hostname only. So let's use the Hostname from
the mgr command output instead.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-28 09:19:01 +02:00
Blaine Gardner f42e4eebfa test: clean up object store tests
In general, make the output from the object store test easier to follow
and debug.

- Make sure assertions are associated with the appropriate subtest.
- Remove `t.Helper()` from complex helpers.
- Add info to log messages.
- Move long parts of (sub-)test names to log messages.
- Move some reused functions back to ceph_base_object_test.go
- Refactor runObjectE2ETestLite() to be suitable for use as the second
  store created in the "run a second object store" test.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-27 12:29:17 -06:00
Sébastien Han 33dbaba38b ci: add daily jobs
We have new jobs now:

* one that runs both smoke and object on the next Pacific version
* one that runs both smoke and object on Ceph master
* one that tests the upgrade from the current pacific stable to the
  pacific devel
* one that tests the upgrade from the current octopus stable to the
  octopus devel

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-27 11:59:04 +02:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Blaine Gardner def4ca57fe Merge pull request #8796 from BlaineEXE/test-multiple-object-stores
test: run multiple object stores at once in CI
2021-09-23 13:31:47 -06:00
Travis Nielsen 09a5b08fde Merge pull request #8766 from thotz/regionfixobcprovisioning
ceph: pass region to newS3agent()
2021-09-23 11:24:55 -06:00
Juan Miguel Olmo Martínez 1649df0462 ceph: retry <ceph orch> commands when they fail
fixes: https://github.com/rook/rook/issues/8759

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2021-09-23 14:29:26 +02:00
Blaine Gardner b9ebf280fa test: run multiple object stores at once in CI
Multiple CephObjectStores should be able to run at the same time. In CI,
test this by:
1. Create the primary object store under test (store A) (this was
   already done)
2. Verify that store A starts and becomes ready (already done but is
   moved here)
3. Create a second object store (store B) (new)
4. Verify that store B starts and becomes ready (new)
5. Delete store B, and ensure it is deleted (new)
6., etc. Continue to verify primary store A operations as before (this
  was already done)

This is necessary to ensure that the Rook operator is using
`radosgw-admin` and the RGW Admin Ops API properly to support multiple
ObjectStores running at once.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-22 13:27:35 -06:00
Travis Nielsen 9ff0753514 test: run all integration tests against the local build
The integration tests must always be run against the local
build of rook, and an image should never be pulled from dockerhub.
To prevent pulling a release or master tag, the local build
will use a tag specific to the build and not ever published
elsewhere.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit a8a40428b0)
2021-09-22 07:39:42 -06:00
Sébastien Han 470fbfd341 Merge pull request #8743 from leseb/next-pacific
ceph: use next ceph v16.2.6 pacific version
2021-09-21 16:41:45 +02:00
Sébastien Han c1a88f34d4 mds: change init sequence
The MDS core team suggested with deploy the MDS daemon first and then do
the filesystem creation and configuration. Reversing the sequence lets
us avoid spurious FS_DOWN warnings when creating the filesystem.

Closes: #8745
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 15:43:34 +02:00
Jiffin Tony Thottan 280c29f330 ceph: pass region to newS3agent()
If the region is specified in the storage class of OBC, use that in the
newS3agent() than using constant "us-east-1".

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-21 13:08:18 +05:30
Juan Miguel Olmo Martínez 5c75139581 ceph: fix probable cause of intermittent fails in the manager test
Explicitly set the length of the string parameter in json.Unmarshal method

fixes: https://github.com/rook/rook/issues/8669

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2021-09-20 13:29:45 +00:00
Blaine Gardner 264256acfb test: remove unused minimal test versions
The `*SuiteMinimalTestVersion` vars in
`tests/integration/ceph_base_deploy_test.go` are no longer needed with
Jenkins no longer being used. Remove them.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 16:49:29 -06:00
Blaine Gardner c60bf241e1 test: make object e2e test its own test
The CephSmokeSuite is becoming quite large and long, and most of the
length is now related to the object e2e tests. Separate the full object
e2e test into CephObjectSuite, and only test the object 'lite' test in
the CephSmokeSuite.

Resolves #8714

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 15:27:27 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Jiffin Tony Thottan 50ecff8f13 ceph: addressing nits from #8211
Addressing remaining nits from the PR #8211

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-09 12:57:32 +05:30
Jiffin Tony Thottan ca43800119 ceph: add options for cephobjectstore user
Adding options for quota, bucket limit, caps for the
`cephobjectstoreuser`.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-07 22:43:09 +05:30
Sébastien Han 8f42bee563 ci: fix object store test
We just need to wait longer when the status is not ready. We needed
another sleep otherwise the status was never nil and the loop went too
fast. See:

```
2021-09-02 16:29:11.372249 I | integrationTest:
2021-09-02 16:29:11.374427 I | integrationTest:
2021-09-02 16:29:11.377764 I | integrationTest:
2021-09-02 16:29:11.379950 I | integrationTest:
2021-09-02 16:29:11.382084 I | integrationTest:
2021-09-02 16:29:11.385383 I | integrationTest:
2021-09-02 16:29:11.388499 I | integrationTest:
2021-09-02 16:29:11.391301 I | integrationTest:
2021-09-02 16:29:11.393545 I | integrationTest:
2021-09-02 16:29:11.396249 I | integrationTest:
```

Signed-off-by: Sébastien Han <seb@redhat.com>

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-07 17:15:20 +02:00
Blaine Gardner a1814af1d9 ceph: remove NFS and Cassandra operator code
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-08-31 14:07:02 -06:00
Travis Nielsen 5489dd7b58 cassandra: suppress integration test errors temporarily
The Cassandra tests are failing in the CI, although they pass when
run locally on a developer cluster. Until the issue can be tracked
down we need to disable these checks to get to a green CI again.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-08-26 14:34:40 -06:00
subhamkrai ebe21c8f68 ci: add action for shellcheck linter
we are adding new linter for shellcheck.
As we are writing more shell scripts this
will help maintain quality.

Also, doing all the changes required in
bash files to pass this shellcheck.

Closes: https://github.com/rook/rook/issues/8431
Signed-off-by: subhamkrai <srai@redhat.com>
2021-08-26 10:03:32 +05:30
Jiffin Tony Thottan f4bb47e440 ceph: add support for update() from lib-bucket-provisioner
Recently lib-bucket-provisioner add support for update() API.
Include that on the obc implementation since it can be used to
update quota for OBC.

Fixes: #7146

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-08-19 23:57:11 +05:30
Juan Miguel Olmo Martínez 6ff5a5b86c ceph: fix error in <ceph orch device ls> test
Rook orchestrator is going to use PVs as <devices> for OSDs.
These PVs must use one storage class present in the cluster and
configured properly in the manager.
With this change we include this new requirement in the Rook manager module test

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2021-08-18 14:19:23 +02:00
parth-gr b28455245d ci: fix for CephObjectStores flakiness
Integration test CephSmokeSuite fails frequently
A quick fix for it by reordering storeName,
running tlsteststore before teststore

Closes: https://github.com/rook/rook/issues/8309
Signed-off-by: parth-gr <paarora@redhat.com>
2021-08-09 14:08:32 +05:30
Juan Miguel Olmo Martínez 9edff582c3 ceph: enable again the Rook orchestrator mgr test
Enable again the mgr test:
- Now is more reliable and robust the start of the test.
- Minor fixes to adapt the <service ls>  to the new name of the crash daemon
deployed by rook

- The creation of OSDs is disabled in the orchestrator, so i have removed
this test until we will have the functionality ready again in the orchestrator
part

My plan is to provide in the orchestrator two different ways to create OSDs:
- Creation of OSD using specific devices if discovery daemon is running
- Creation of OSds using PVS (if we have LSO/other LS operator running)

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2021-07-28 16:27:42 +02:00
Travis Nielsen d1f02f22f3 ceph: test upgrades from v1.6 to master
With v1.7 approaching, the upgrade integration test will now test from
v1.6.x to the latest master, which will effectively become the v1.7
release soon.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-20 11:03:40 -06:00
Sébastien Han eaa6e7732c Merge pull request #8272 from leseb/exec-in-pod
ceph: proxy ceph command when multus is configured
2021-07-07 21:36:54 +02:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Sébastien Han 830b36c039 Merge pull request #7920 from thotz/tlscitest
test: ci test for TLS objectstore
2021-07-05 11:23:40 +02:00
Sébastien Han baaea4a1ea Merge pull request #7604 from leseb/cephfs-mirror-peer-config
ceph: add filesystem mirror peers configuration
2021-07-05 10:36:30 +02:00
Travis Nielsen 7cec49b2d8 ceph: run helm tests without admission controller
The helm tests frequently hit an issue with the admission controller.
The admission controller is also enabled in other test suites, so
it seems safe enough to disable for the helm test suite.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-02 15:28:57 -06:00
Sébastien Han b578f916e7 ceph: add fs mirror config
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.

So the automatic configuration of Ceph Filesystem peers is now possible.

By editing the CephFilesystem CRD, you can now turn on mirroring:

```yaml
  mirroring:
    enabled: false
    # list of Kubernetes Secrets containing the peer token
    # for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
    peers:
      secretNames:
        - secondary-cluster-peer
```

Also, the mirroring status is displayed in the CR status:

```
status:
  info:
    fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
  mirroringStatus:
    daemonsStatus:
    - daemon_id: 4186
      filesystems:
      - filesystem_id: 2
        name: myfs
    lastChecked: "2021-07-01T14:16:29Z"
  phase: Ready
  snapshotScheduleStatus:
    lastChecked: "2021-07-01T14:16:29Z"
    snapshotSchedules:
    - fs: myfs
      path: /
      rel_path: /
      retention: {}
      schedule: 24h
```

Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 17:35:19 +02:00
Jiffin Tony Thottan 9e3cf68d04 test: ci test for TLS objectstore
Extend the object store smoke test to include TLS configurations.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-07-01 18:50:27 +05:30