Commit Graph
283 Commits
Author SHA1 Message Date
Travis Nielsen fc541764f1 Merge pull request #10484 from jsoref/spelling
core: fix spelling
2022-07-15 08:31:41 -06:00
Travis Nielsen 53955790f3 ci: pin the whereabouts version to v0.5.3
The whereabouts manifests in the master branch
have moved around, so until the new approach
is investigated we pin to the latest release
version v0.5.3

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-07-13 10:37:15 -06:00
Satoru Takeuchi 7e571f6114 osd: support OSD on logical volume in host-based cluster
Rook supports raw mode OSD in host-based cluster. So we can also
support OSD on logical volume in this kind of cluster.

Logical volumes aren't picked by filters (i.e. `useAllDevices: true`
and `device{Path,}Filter` to avoid unwanted LV consumption on upgrade.

Closes: https://github.com/rook/rook/issues/2047

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-07-11 21:08:50 +00:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Satoru Takeuchi 6b0f539f61 test: add canary test for encrypted osd in host-based clusters
It's better to test the creation of encrypted devices in host-based clusters.
This test will reduce the potential risks of regression when modifying
the osd-creation code.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-06-08 22:34:01 +00:00
Travis Nielsen 1acc0621af Merge pull request #10342 from Madhu-1/fix-csi-pod-network
ceph: Add holder pod if csi host networking is disabled
2022-06-07 10:16:51 -06:00
subhamkrai d5bac59a74 ci: use python formatting tool black
formatting other python files using black

Signed-off-by: subhamkrai <srai@redhat.com>
2022-06-07 13:49:14 +05:30
Madhu Rajanna 276005cfd5 csi: add holder pod if csi hostnetworking is disabled
If csi is configured not the use the
hostnetworking, deploy the holder
pod for executing the commands with nsenter.
The implementation is same as multus.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-06-07 13:09:12 +05:30
Satoru Takeuchi 5b3633bfd6 test: add canary integration test for osd with metadata device
"metadataDevice" field in host based cluster is not tested yet.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-05-25 01:36:33 +00:00
Satoru Takeuchi a9465efa07 Merge pull request #9931 from cybozu-go/test-add-canary-integration-tests-for-osd-on-device
test: add canary integration tests for osd on device
2022-05-06 05:25:57 +09:00
Satoru Takeuchi 805163e2e2 test: add canary integration tests for osd on device
OSD on PVC is tested in many patterns but OSD on device is not.
It's preferable to add the following patterns for OSD on device.

- An OSD on a raw disk
- Two OSDs on a device

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-05-05 13:15:18 +00:00
Satoru Takeuchi ab38c10c2d test: avoid to create a corrupted GPT headers
The following steps in CI scripts create corrupted GPT headers.

```console
$ sudo sgdisk --zap-all --clear --mbrtogpt -g -- "$DISK"
$ sudo dd if=/dev/zero of="$DISK" bs=1M count=10
```

It results in the failure of the succeeding `sgdisk --print "$DISK".

Here is an example.

```console
$ sudo parted /dev/sdb mklabel msdos
...
$ sudo sgdisk --zap-all --clear --mbrtogpt -g -- /dev/sdb
...
$ sudo dd if=/dev/zero of=/dev/sdb bs=1M count=10
...
$ sudo sgdisk --print /dev/sdb
Caution: invalid main GPT header, but valid backup; regenerating main header
from backup!

Warning: Invalid CRC on main header data; loaded backup partition table.
Warning! One or more CRCs don't match. You should repair the disk!
Main header: ERROR
Backup header: OK
Main partition table: OK
Backup partition table: OK

Invalid partition data!
$ echo $?
2
```

We can safely use the simple `sgdisk --zap-all "$DEVICE"` here.
It's not necessary to convert "mbr" to "gpt" because we'll make new GPT
labels just after this command.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-05-05 13:11:21 +00:00
Sébastien Han 6ef491d9fc osd: close the encrypted disk after cleanup is done
Once we are done cleaning up the content of the encrypted osd, let's
close the main LUKS device.
This avoids having dm devices on the system after the cleanup.

Closes: https://github.com/rook/rook/issues/10181
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-05-03 16:47:59 +02:00
Sébastien Han 4ea8cc6224 Merge pull request #10173 from subhamkrai/fix-deploy-vault
ci: wait for cert while deploying vault
2022-04-27 11:51:02 +02:00
Sébastien Han ed3121defd Merge pull request #9925 from leseb/multus-plugin-restart-fix
core: fix csi-cephfsplugin pod restart on non-hostnetworking env
2022-04-27 11:41:59 +02:00
subhamkrai 86b3d466f6 ci: wait for cert while deploying vault
sometimes, it take few seconds to fill certificate
after certificate approval. So, let's wait for few seconds.

Closes: https://github.com/rook/rook/issues/10158
Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-27 14:44:24 +05:30
Sébastien Han 73b1347675 core: fix csi-cephfsplugin pod restart on non-hostnetworking env
Implementation of the design proposed in
https://github.com/rook/rook/pull/9903.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 17:53:01 +02:00
Sébastien Han 79dbb35302 ci: collect secrets and cms in the logs
The CI runs will now collect all the secrets and configmaps available.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 11:05:02 +02:00
Travis Nielsen d26b6cf403 build: update min version to k8s 1.17
Support is removed for k8s for various limitations such as
priority classes not working and csi driver feature
incompleteness and missing snapshots. Documentation is updated
with the new min version of k8s 1.17 and also the tests are
updated to run on the min version of 1.17.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-25 15:32:09 -06:00
Madhu Rajanna 4425ef1ce2 build: provide 775 permission for deploy_cert_manager.sh
provided 775 permission like other scripts for
deploy_cert_manager.sh

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-04-13 10:57:33 +05:30
Sébastien Han 2c09bbb91b ci: add multus integration test
This new integration test will deploy a cluster with multus enabled. It
will be comprised of two network interfaces for ceph public and cluster
communications.
For now, it only deploys a Ceph cluster up to the OSDs.

Closes: https://github.com/rook/rook/issues/9784
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-12 17:30:47 +02:00
subhamkrai f6f03d272b core: start admission controller without any script
finally, admission controller will be enabled default
without any script/manual step. But it still requires cert-manager
to be installed which I believe is already installed in clusters.

**Note**
Code doesn't return error it just logs the error since
we don't want to stop reconciling if the admission controller fails.
We can work on this once the admission controller is stable.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-11 19:44:54 +05:30
Travis Nielsen 4f9a72c1db Merge pull request #9894 from yuvalman/resources
helm: enable resource defaults on rook components
2022-03-21 13:58:17 -06:00
Yuval Manor ff7a5c2d2d helm: enable resource defaults on rook components
This PR assign default values to the resources of all Rook and Ceph components

Closes: https://github.com/rook/rook/issues/9858
Signed-off-by: Yuval Manor <yuvalman958@gmail.com>
2022-03-21 21:23:35 +02:00
Satoru Takeuchi 3bbffd1e4d test: avoid potential data inconsistency on zapping disk
Data inconsistency might happen if the disk is accessed just after disk
zapping because direct I/O is not synchronous by itself.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-03-18 11:50:04 +00:00
Sébastien Han 9fcc6891b6 Merge pull request #9706 from subhamkrai/use-webhook-placeholder-value
core: remove caBundle value from webhook
2022-03-14 08:59:15 +01:00
subhamkrai 2115daa317 core: refactor webhook config and related script
This commit refactor webhook config by adding/removing
spaces. Also, moved the Issue and Certificate creation
from bash script to webhook config yaml.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-03-14 10:16:01 +05:30
Travis Nielsen 75a385e187 csi: update volume replication to v0.3.0
Update to the latest volume replication crd version of
v0.3.0

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-03-09 15:53:45 -07:00
subhamkrai 536a73af64 core: remove caBundle value from webhook
now, we don't require caBundle values to verify in
webhook config.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-03-09 09:11:02 +05:30
Blaine Gardner 9a5d43761d test: collect fuller logs for CI tests
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-02-22 13:14:39 -07:00
Sébastien Han d86351aa1c Merge pull request #9728 from thotz/vault-kms-backendpath-tranist-engine-fix
object: fix backend path for transit engine for rgw kms
2022-02-16 10:21:08 +01:00
Jiffin Tony Thottan 0b4cdb992b object: fix backend path for transit engine for rgw kms
The backend path was added with additional `transit` to it.
Also added PR test case to check transit engine

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-02-15 18:49:31 +05:30
Jiffin Tony Thottan 438cf6abf1 test: update bucket notification integration test with http server
Add test cases to check notification is received by http server. For
this a sample http server https://github.com/thotz/pythonwebserver is
used.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-02-15 12:59:58 +05:30
Travis Nielsen 50afb9f026 helm: update to the latest helm version v3.8
The CI was building with helm 3.6.2, now updating to
the latest v3.8.0

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-01-24 14:12:22 -07:00
Blaine Gardner bb0d0d39f6 Merge pull request #9384 from leseb/fix-7036
subvolumegroup: add new crd
2021-12-21 10:24:23 -07:00
Sébastien Han 6e9eb33782 subvolumegroup: add new crd
This introduces a new CRD to add the ability to create subvolumegroup
for a given ceph filesystem volume. Typically the name of the volume is
the name of the filesystem created by rook.

Closes: https://github.com/rook/rook/issues/7036
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-21 16:28:38 +01:00
Sébastien Han 05775b068d osd: handle removal of encrypted osd deployment
This is handling a tricky scenario where the OSD deployment is manually
removed and the OSD never reconvers. This is unlikely to happen, but
still OSD should be able to run after that action. Essentially after a
manual deletion, we need to run the prepare job again to re-hydrate the
OSD information so that the OSD deployment can be deployed.
On encryption, it is a little bit tricky since ceph-volume list again
the main block won't return anything, so we need to target the encrypted
block to list.
There is another case this PR does not handle, which is the removal of
the OSD deployment and then the node is restarted. This means that the
encrypted container is not opened anymore. However, opening it requires
more work like writing the key on the filesystem (if not coming from the
Kubernete secret, eg,. KMS vault) and then run luksOpen. This is an
extreme corner case probably not worth worrying about for now.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-21 10:51:09 +01:00
Travis Nielsen 60cc67fed4 test: only create replicated pools in ci
Creating an EC pool is causing the CI to hang when the EC
pool is initialized since there aren't enough OSDs to
satisfy the EC parameters.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-15 15:32:03 -07:00
subhamkrai 22865815cf pool: add rbd ec pool support in external cluster
Adding cli agrument `--rbd-metadata-ec-pool-name` to read
rbd ec pool name to support ec pool in external cluster
and also updating the json blob.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-12-14 12:13:22 +05:30
Blaine Gardner 243a47bfc4 build: do not use cross build container for ceph
Do not use the cross build container when building, publishing, and
promoting rook/ceph images. It is no longer needed, and its complexity
can add flakiness.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-03 11:06:21 -07:00
Sébastien Han c890710b63 core: change directory layout
As per discussion, proposing a new layout for the charts/yaml/olm files.

./deploy
├── charts
│   ├── rook-ceph
│   │   └── templates
│   └── rook-ceph-cluster
│       └── templates
├── examples
│   ├── csi
│   │   ├── cephfs
│   │   └── rbd
│   ├── flex
│   ├── monitoring
│   ├── pre-k8s-1.16
└── olm
    └── assemble

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-30 09:12:53 +01:00
Sébastien Han 7a223adde4 Merge pull request #9230 from leseb/osd-rm-check
osd: check if osd is safe-to-destroy before removal
2021-11-25 10:55:34 +01:00
Sébastien Han ad2c3c2ae6 Merge pull request #8931 from BlaineEXE/test-rgw-multisite-in-nightly-tests
test: test rgw multisite nightly
2021-11-25 10:25:57 +01:00
Sébastien Han 7402c2cce6 osd: check if osd is ok-to-stop before removal
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.

So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-25 10:21:32 +01:00
Travis Nielsen d340bccd44 Merge pull request #9202 from thotz/hpakedadoc
docs: add details about HPA via KEDA
2021-11-23 10:42:03 -07:00
Jiffin Tony Thottan d32c837e35 docs: add details about HPA via KEDA
By default HPA can use details about memory or CPU consumption for
autoscaling, but also it can use custom metrics as well. There are alot
provides supports HPA via customer and one of them is KEDA project. Here
it is done with help of Prometheus Scaler

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-11-23 16:50:18 +05:30
Blaine Gardner e80368674d test: run RGW multisite test in nightly job
In the nightly job, run the test with the latest Ceph version so we can
detect if there are RGW changes in Ceph that might break multisite. Use
a reusable GitHub action workflow to duplicate as little code as
possible.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-19 09:56:41 -07:00
Jiffin Tony Thottan aba50d3ca9 object: add support in RGW to communicate vault with TLS
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-11-17 10:19:28 +05:30
Sébastien Han 0e26176b54 ci: wait for kubeproxy to be ready
When requesting issuer, we run a local kubectl proxy command which spawn
a proxy server. However, we must wait for the proxy to be ready before
we actually start making requests to it.
Now the CI waits up to 10sec to retrieve the issuer.

Closes: https://github.com/rook/rook/issues/9090
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-08 17:45:53 +01:00
Sébastien Han 16729e0e38 osd: use multiple service account for vault role
We can pass bound_service_account_names with a comma separated list of
service accounts. Let's do this instead of remapping new values.
Earlier, we thought a single service account could be added per Vault
role and we were using other variables like
`VAULT_AUTH_KUBERNETES_ROOK_OPERATOR_ROLE` that we were remapping to
`VAULT_AUTH_KUBERNETES_ROLE` internal for the API calls to Vault.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-04 15:15:46 +01:00