Commit Graph
33 Commits
Author SHA1 Message Date
Blaine Gardner 5539eedd1b multus: finish deprecating holder pods
Finish the process of deprecating holder pods by removing Rook's ability
to deploy them. The intent of this change is to make the most
superficial changes possible to accomplish this. There are still
remnants of code in Rook (particularly the CSI controller) that helped
configure or deploy holder pods. Due to the risk of breaking some
features, cleanup work of hose remnants will be deferred for future
work.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-10-23 16:29:02 -06:00
Travis Nielsen fdacfd51c5 object: create an object store based on shared pools
Until now, an object store would create all the necessary
metadata pools and the data pool that were exclusively
for its own object store. When isolation between object
stores is necessary, this would cause many pools and
PGs to be created in the cluster, which was not
manageable.

Now one set of pools can be created to be shared
by any number of object stores. The metadata and data
between each object store is isolated by
RADOS namespaces, which by design will keep the
data safe for multi-tenancy.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-03-11 11:20:57 -06:00
Blaine Gardner 3e2e906ea7 test: update multus canary test
Update the multus canary test to reflect modern knowledge about how it
should be configured.

No longer test for the network device in OSD pods. Pods will utterly
fail to start if Multus is unable to attach interfaces.
Instead, look to the OSD map to test the connections more wholistically.
OSDs must have map IPs that include both public and cluster network.
This implicitly tests that the interfaces exist in the Pod, and it
additionally verifies other details, like Ceph `*_network` configs are
set propertly.

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2023-09-18 17:15:23 -06:00
subhamkrai 4502ea8ee1 multus: use right interface in ci validation
runner version `2.306` had the interface `net`
but somehow version `2.307.1` which is latest
doesn't have `net` it has `eth0*` so using that.

```
cat /proc/net/dev
Inter-|   Receive                                                |  Transmit
 face |bytes    packets errs drop fifo frame compressed multicast|bytes    packets errs drop fifo colls carrier compressed
    lo:       0       0    0    0    0     0          0         0        0       0    0    0    0     0       0          0
 tunl0:       0       0    0    0    0     0          0         0        0       0    0    0    0     0       0          0
  eth0:     446       5    0    0    0     0          0         0        0       0    0    0    0     0       0          0
sh-4.4# grep etho /proc/net/dev
sh-4.4# grep eth0 /proc/net/dev
  eth0:     446       5    0    0    0     0          0         0        0       0    0    0    0     0       0          0
sh-4.4# exit
```

Signed-off-by: subhamkrai <srai@redhat.com>
2023-08-02 20:11:34 +05:30
Jiffin Tony Thottan 715c89bcd8 test: check rgw pod is running for canary github workflow
For the RGW daemon validation please check whether pod is Running than
the exisitng checks

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-04-06 16:28:10 +05:30
Travis Nielsen 4c5163f843 test: print return value for intermittent failure
The canary tests sometimes have an intermittent failure
when processing the return value of a grep for the osds
to be running. Print some debug info to help track it
down.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-23 09:11:20 -07:00
Travis Nielsen ff01722b0f test: increase timeout waiting for ceph commands
The test scripts were only waiting for a timeout of three
seconds for ceph commands, which was causing intermittent
failures in the CI. Now the timeout is increased to
ten seconds.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2023-02-23 09:11:20 -07:00
Rakshith R 07aac106df ci: add e2e for csi-nfsplugin restart
This commit adds e2e for csi-nfsplugin restart
when it is not on host-networking.

Signed-off-by: Rakshith R <rar@redhat.com>
2022-09-02 14:16:15 +05:30
Madhu Rajanna 276005cfd5 csi: add holder pod if csi hostnetworking is disabled
If csi is configured not the use the
hostnetworking, deploy the holder
pod for executing the commands with nsenter.
The implementation is same as multus.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-06-07 13:09:12 +05:30
Sébastien Han 73b1347675 core: fix csi-cephfsplugin pod restart on non-hostnetworking env
Implementation of the design proposed in
https://github.com/rook/rook/pull/9903.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 17:53:01 +02:00
Sébastien Han 2c09bbb91b ci: add multus integration test
This new integration test will deploy a cluster with multus enabled. It
will be comprised of two network interfaces for ceph public and cluster
communications.
For now, it only deploys a Ceph cluster up to the OSDs.

Closes: https://github.com/rook/rook/issues/9784
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-12 17:30:47 +02:00
Sébastien Han e0145b9643 nfs: add pool setting CR option
Ths NFS spec now supports the CephBlockPool spec which means that it can
take advantage of all the known settings like compression, size, failure
domain etc.

Closes: https://github.com/rook/rook/issues/9034
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-29 09:04:31 +02:00
subhamkrai e49cdf0a8d ci: clean validate_cluster.sh script
removing `trap display_status SIGINT ERR` command
from the file as `display_status` func has been removed.

Closes: https://github.com/rook/rook/issues/9004
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 10:43:50 +05:30
Blaine Gardner 65619c2677 test: get more partition info setting up ci disk
Add some commands to get more partition info when setting up the GH
action runner's disk for use in integration tests. This will both aid in
debugging and may "jog" the system such that it will no longer need to
reload the partition info when running the OSD prepare job.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-18 11:29:13 -06:00
Blaine Gardner f287df70ee test: fix prepare pod log collection in CI
In the CI tests that use `validate_cluster.sh display_status` to gather
logs, the prepare pod log collection failed. Fix this.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-14 11:15:59 -06:00
Blaine Gardner 34a8b42e97 test: try to un-flake multi-cluster-mirror test
Try to un-flake the multi-cluster-mirror test that keeps failing on this
PR.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-13 11:06:51 -06:00
Blaine Gardner a1dd256d4b Merge pull request #8911 from BlaineEXE/rgw-commands-use-staging-flag
rgw: replace period update --commit with function
2021-10-12 08:22:41 -06:00
Blaine Gardner 956430826c rgw: add integration test for committing period
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.

Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.

To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
   of its operation that didn't cause failures in the past

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-10-11 15:24:59 -06:00
Sébastien Han 4f9c31fa8b ci: clarify the wait for csi to be ready
The wait is now more comprehensive.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-11 16:47:02 +02:00
Sébastien Han 85216a266c ci: wait longer for csi to be available
Sometimes the CI needs more time...

Closes: https://github.com/rook/rook/issues/8825
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-27 12:12:46 +02:00
Sébastien Han 73c340ded7 ci: fix pod list
We should not use .items[0].metadata.name if the array length is 0. This
is the case when nothing has been initialized yet. Instead, we should
use .items[*].metadata.name, the wildcard ensures to always return 0
even if nothing is present yet.

Fixes: https://github.com/rook/rook/issues/8676
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-09 17:24:56 +02:00
Sébastien Han b578f916e7 ceph: add fs mirror config
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.

So the automatic configuration of Ceph Filesystem peers is now possible.

By editing the CephFilesystem CRD, you can now turn on mirroring:

```yaml
  mirroring:
    enabled: false
    # list of Kubernetes Secrets containing the peer token
    # for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
    peers:
      secretNames:
        - secondary-cluster-peer
```

Also, the mirroring status is displayed in the CR status:

```
status:
  info:
    fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
  mirroringStatus:
    daemonsStatus:
    - daemon_id: 4186
      filesystems:
      - filesystem_id: 2
        name: myfs
    lastChecked: "2021-07-01T14:16:29Z"
  phase: Ready
  snapshotScheduleStatus:
    lastChecked: "2021-07-01T14:16:29Z"
    snapshotSchedules:
    - fs: myfs
      path: /
      rel_path: /
      retention: {}
      schedule: 24h
```

Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 17:35:19 +02:00
Sébastien Han 7c42b61f3c ci: add more debug logs to failed job
Gather more logs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-31 14:16:46 +02:00
Travis Nielsen 665856b90a ceph: enable pacific as a supported ceph version
With the Ceph Pacific release coming this week we add support
in Rook for Pacific with the Rook v1.6 release coming soon.
The integration tests will now run across nautilus, octopus,
and pacific to cover all supported Ceph versions. The default
examples still specify Octopus until there is more bake time
for Pacific.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-04-07 06:59:04 -06:00
Travis Nielsen f620df9fd8 ceph: remove duplicate function to validate rgw in the ci
The CI had a duplicate function for validating the rgw daemon
was running

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-04-07 06:59:04 -06:00
Sébastien Han cc822a531c ceph: add more osds to the ci
We now create 3 disks which will give us 2 OSDs and potentially help us
to catch more errors in our scenarios, especially on iterations.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-09 15:24:35 +01:00
Sébastien Han c0123cf182 ceph: add cephfs mirroring support
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.

The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050

The same goes for integration tests, they will get added later once we
start testing on Pacific.

Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-28 19:21:18 +01:00
subhamkrai fcd1529b23 ci: add option to download artifacts in canary test
this commit add option to download artifacts in
canary integration test for better debug.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-01-12 21:06:30 +05:30
subhamkrai f2ea2daf07 ceph: remove extra space
this commit removes extra space from
func test_csi.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-01-05 14:58:35 +05:30
subhamkrai 6ea06713b6 ceph: remove leftover line
this commit removes the leftover line from PR #6800.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-01-05 10:46:51 +05:30
subhamkrai 24c3002307 ci: minor update to validate script
canary test failing because the validate
script not updating the argument passed.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-12-10 13:12:39 +05:30
Sébastien Han 3403776c3e ci: add more osd on pvc scenario
Add the following scenario:

* simple osd on pvc
* osd on pvc with db device
* osd on pvc with wal device
* encrypted simple osd on pvc
* encrypted osd on pvc with db device
* encrypted osd on pvc with wal device

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-30 16:17:10 +01:00
Sébastien Han ea1d71cbfb ceph: add vault kms support for osd encryption
When the Ceph cluster runs on PVC and the OSDs are encrypted we can
store LUKS's Key Encryption Key inside a Key Management System. Today,
Rook only supports HashiCorp Vault: https://www.vaultproject.io/

The CephCluster has now a new "security" field which will plug onto the
KMS. Here is an example:

security:
  kms:
    tokenSecretName: <name of the secret containing a Vault token, used
    to authenticate>
    connectionDetails: < a map of strings containing connection
    information>

Refer to the ceph-cluster-crd documentation to lear more.

Closes: https://github.com/rook/rook/issues/6105
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-30 16:16:33 +01:00