Finish the process of deprecating holder pods by removing Rook's ability
to deploy them. The intent of this change is to make the most
superficial changes possible to accomplish this. There are still
remnants of code in Rook (particularly the CSI controller) that helped
configure or deploy holder pods. Due to the risk of breaking some
features, cleanup work of hose remnants will be deferred for future
work.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Until now, an object store would create all the necessary
metadata pools and the data pool that were exclusively
for its own object store. When isolation between object
stores is necessary, this would cause many pools and
PGs to be created in the cluster, which was not
manageable.
Now one set of pools can be created to be shared
by any number of object stores. The metadata and data
between each object store is isolated by
RADOS namespaces, which by design will keep the
data safe for multi-tenancy.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Update the multus canary test to reflect modern knowledge about how it
should be configured.
No longer test for the network device in OSD pods. Pods will utterly
fail to start if Multus is unable to attach interfaces.
Instead, look to the OSD map to test the connections more wholistically.
OSDs must have map IPs that include both public and cluster network.
This implicitly tests that the interfaces exist in the Pod, and it
additionally verifies other details, like Ceph `*_network` configs are
set propertly.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
For the RGW daemon validation please check whether pod is Running than
the exisitng checks
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
The canary tests sometimes have an intermittent failure
when processing the return value of a grep for the osds
to be running. Print some debug info to help track it
down.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The test scripts were only waiting for a timeout of three
seconds for ceph commands, which was causing intermittent
failures in the CI. Now the timeout is increased to
ten seconds.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If csi is configured not the use the
hostnetworking, deploy the holder
pod for executing the commands with nsenter.
The implementation is same as multus.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
This new integration test will deploy a cluster with multus enabled. It
will be comprised of two network interfaces for ceph public and cluster
communications.
For now, it only deploys a Ceph cluster up to the OSDs.
Closes: https://github.com/rook/rook/issues/9784
Signed-off-by: Sébastien Han <seb@redhat.com>
Ths NFS spec now supports the CephBlockPool spec which means that it can
take advantage of all the known settings like compression, size, failure
domain etc.
Closes: https://github.com/rook/rook/issues/9034
Signed-off-by: Sébastien Han <seb@redhat.com>
Add some commands to get more partition info when setting up the GH
action runner's disk for use in integration tests. This will both aid in
debugging and may "jog" the system such that it will no longer need to
reload the partition info when running the OSD prepare job.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
In the CI tests that use `validate_cluster.sh display_status` to gather
logs, the prepare pod log collection failed. Fix this.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.
Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.
To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
of its operation that didn't cause failures in the past
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
We should not use .items[0].metadata.name if the array length is 0. This
is the case when nothing has been initialized yet. Instead, we should
use .items[*].metadata.name, the wildcard ensures to always return 0
even if nothing is present yet.
Fixes: https://github.com/rook/rook/issues/8676
Signed-off-by: Sébastien Han <seb@redhat.com>
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.
So the automatic configuration of Ceph Filesystem peers is now possible.
By editing the CephFilesystem CRD, you can now turn on mirroring:
```yaml
mirroring:
enabled: false
# list of Kubernetes Secrets containing the peer token
# for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
peers:
secretNames:
- secondary-cluster-peer
```
Also, the mirroring status is displayed in the CR status:
```
status:
info:
fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
mirroringStatus:
daemonsStatus:
- daemon_id: 4186
filesystems:
- filesystem_id: 2
name: myfs
lastChecked: "2021-07-01T14:16:29Z"
phase: Ready
snapshotScheduleStatus:
lastChecked: "2021-07-01T14:16:29Z"
snapshotSchedules:
- fs: myfs
path: /
rel_path: /
retention: {}
schedule: 24h
```
Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
With the Ceph Pacific release coming this week we add support
in Rook for Pacific with the Rook v1.6 release coming soon.
The integration tests will now run across nautilus, octopus,
and pacific to cover all supported Ceph versions. The default
examples still specify Octopus until there is more bake time
for Pacific.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We now create 3 disks which will give us 2 OSDs and potentially help us
to catch more errors in our scenarios, especially on iterations.
Signed-off-by: Sébastien Han <seb@redhat.com>
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.
The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050
The same goes for integration tests, they will get added later once we
start testing on Pacific.
Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>
Add the following scenario:
* simple osd on pvc
* osd on pvc with db device
* osd on pvc with wal device
* encrypted simple osd on pvc
* encrypted osd on pvc with db device
* encrypted osd on pvc with wal device
Signed-off-by: Sébastien Han <seb@redhat.com>
When the Ceph cluster runs on PVC and the OSDs are encrypted we can
store LUKS's Key Encryption Key inside a Key Management System. Today,
Rook only supports HashiCorp Vault: https://www.vaultproject.io/
The CephCluster has now a new "security" field which will plug onto the
KMS. Here is an example:
security:
kms:
tokenSecretName: <name of the secret containing a Vault token, used
to authenticate>
connectionDetails: < a map of strings containing connection
information>
Refer to the ceph-cluster-crd documentation to lear more.
Closes: https://github.com/rook/rook/issues/6105
Signed-off-by: Sébastien Han <seb@redhat.com>