Commit Graph
7961 Commits
Author SHA1 Message Date
Travis Nielsen 9d2aa1f6bd test: generate long node name depending on test suite
The generation of a long node name in the integration tests was
being done based on the k8s version. In the past, older K8s versions
did not support the changing name. Now it's more maintainable if
we generate the long name depending on the test suite.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-06 08:03:56 -07:00
Travis Nielsen ee83ca74af osd: truncate osd prepare job names further
In K8s 1.22 there is a bug in the job name generation that
the job name is truncated an additional 10 characters. This can cause an issue
in the generated pod name if it then ends in a non-alphanumeric character. In that case,
we more aggressively generate a hashed job name.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-06 08:02:04 -07:00
Travis Nielsen 117229a6b7 Merge pull request #9282 from travisn/failover-arbiter
mon: Set stretch tiebreaker reliably during failover
2021-12-03 11:28:42 -07:00
Blaine Gardner fc7740489a Merge pull request #9304 from BlaineEXE/debug-ceph-makefile
build: do not use cross build container for ceph
2021-12-03 11:12:58 -07:00
Blaine Gardner 243a47bfc4 build: do not use cross build container for ceph
Do not use the cross build container when building, publishing, and
promoting rook/ceph images. It is no longer needed, and its complexity
can add flakiness.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-03 11:06:21 -07:00
Blaine Gardner 130252e8ce Merge pull request #9279 from BlaineEXE/sort-common.yaml
docs: sort common.yaml
2021-12-03 09:27:19 -07:00
Travis Nielsen 2efa571c93 Merge pull request #9280 from TomHellier/test-helm-on-multiple-versions
helm: Use a more explicit value for ingress for K8s 1.18
2021-12-02 14:30:12 -07:00
Travis Nielsen 8c60f6f5f9 Merge pull request #9292 from travisn/backport-mergify-1.8
bot: Enable backports to release-1.8
2021-12-02 14:29:45 -07:00
Blaine Gardner aba199fc4c docs: sort common.yaml example
In anticipation of generating common.yaml from the Helm charts, sort
common.yaml using the same script used to sort the output from Helm
charts.

This was done using the flow here:
```
cat build/rbac/common.yaml.header > new-common.yaml
cat deploy/examples/common.yaml | build/rbac/keep-rbac-yaml.sh >> new-common.yaml
mv new-common.yaml deploy/examples/common.yaml
```

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-02 14:24:14 -07:00
Travis Nielsen d55669293a mon: set stretch tiebreaker reliably during failover
The failover of the arbiter mon in a stretch cluster was sometimes
failing due to the new tiebreaker not being set in ceph.
Rook would repeatedly try to remove the old tiebreaker mon
and keep failing because the new tiebreaker had not been set.
Now we make setting the tiebreaker idempotent in case the operator
restarts in the middle of the operation or some other corner
case causes the expected tiebreaker to be set. In that case,
the next reconcile will also ensure the tiebreaker mon is
set as expected.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-02 12:50:07 -07:00
Travis Nielsen f7b9ca966e bot: enable backports to release-1.8
With the creation of the release-1.8 branch we enable
the mergify bot to open the backport PRs automatically
based on the backport-release-1.8 label.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-02 11:51:57 -07:00
Blaine Gardner 67e6220e85 Merge pull request #9200 from BlaineEXE/generate-csv-without-comments
build: generate CSV from helm charts
2021-12-02 10:42:50 -07:00
Satoru Takeuchi bdbb97470d Merge pull request #9290 from travisn/skip-osd-checks
osd: Honor skipUpgradeChecks for OSDs
2021-12-02 09:47:07 +09:00
Tom Hellier 2ca93e547c helm: use a more explicit value for ingress
The networking.k8s.io/v1 api defines the backend differently
to networking.k8s.io/v1beta1 and extensions/v1beta1

Signed-off-by: Tom Hellier <me@tomhellier.com>
2021-12-01 23:39:52 +00:00
Travis Nielsen 32a884ac18 osd: honor skipUpgradeChecks for osds
Skipping upgrade checks was not being honored for OSDs.
Now the flag will be checked and allow the OSDs to be upgraded
without checking for the ok-to-stop condition.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-01 14:41:53 -07:00
Travis Nielsen baa67c8d5a Merge pull request #9287 from TomHellier/9286-add-mount-options-to-helm-chart
helm: addition of mountOptions into storage class configuration
v1.8.0-alpha.0
2021-12-01 10:19:35 -07:00
Sébastien Han 70fab7c042 Merge pull request #9288 from leseb/add-missing-crds
build: add missing topic and notification crd to csv
2021-12-01 18:08:21 +01:00
Travis Nielsen a700ecbfe1 Merge pull request #8547 from humblec/selinux
csi: mount host's /etc/selinux in node plugins
2021-12-01 08:59:45 -07:00
Sébastien Han b4a8b01c0e build: add missing topic and notification crd to csv
Add the missing CRDs as well as a check to prevent missing CRDs from the
CSV file.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-01 16:48:42 +01:00
Tom Hellier caa2c8323b helm: addition of mountOptions into storage class configuration
It should be possible to configure the storage classs mount options, this follows
the helm code used by the ceph-csi project for their ceph-csi-rbd and ceph-csi-cephfs
helm charts.

Signed-off-by: Tom Hellier <me@tomhellier.com>
2021-12-01 10:50:42 +00:00
Humble Chirammal e5f5d9be9f csi: mount host's /etc/selinux in node plugins
This commit introduces a new configuration option for
ceph csi driver to enable hostpath mounting of /etc/selinux
directory from the cluster node where csi plugin pods are
running, which inturn help the csi driver to specify
selinux-related mount options like context.

Ref# https://github.com/ceph/ceph-csi/issues/2295

The default value for this configuration is true and if cluster
nodes are running without selinux enabled, an admin can deploy
csi pods by specifying this option to `false` which skip the
host path mounting for the csi pods.

Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
2021-12-01 11:13:26 +05:30
Sébastien Han 4ec496dd47 Merge pull request #9273 from travisn/upgrade-test-from-1.7
test: Upgrade integration test from 1.7 to master
2021-11-30 17:05:49 +01:00
Blaine Gardner 2c4b66957c build: generate CSV from helm charts
In order to generate common.yaml from Helm charts, we have to have a
different way of generating CSV than from meta-comments in common.yaml.
This implementation changes what appears in the CSV's RBAC somewhat, but
these changes could be considered bugs fixed by using the new Helm
generation method.
- a few PSP related resources are removed from CSV
- resources related to 'rook-ceph-purge-osd' Job are added to CSV
- ClusterRoleBinding 'rook-ceph-object-bucket' is added to CSV

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-30 08:35:54 -07:00
Blaine Gardner 5e784a73ac Merge pull request #9223 from BlaineEXE/use-yq-v4
build: use yq for RBAC yaml parsing
2021-11-30 08:35:30 -07:00
Sébastien Han 4a5ba05777 Merge pull request #9272 from olivierbouffet/fix9100
object: fix search user in objectstore
2021-11-30 16:23:30 +01:00
Blaine Gardner 51fc36d591 build: prevent parallel build collisions for ceph
Prevent parallel build collisions in the 'ceph' image by building
prerequisites for the builds before running the build targets in
parallel. Build targets can collide and try to create the same target at
nearly the same time, causing failures.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-30 08:09:24 -07:00
Blaine Gardner 113e2f849d build: use yq for RBAC yaml parsing
Use yq instead of Python for parsing RBAC from the Helm chart. We need
to use yq v4.14.1 or higher to fix yq's handling of the yaml header
markers ('---'). Update the Makefile's yq version to v4, which also
requires updating the script to update the CRDs. This was quite easy.

It is very difficult, however, to change the version of yq used by the
CSV generating/parsing scripts, which already used their own yq
download. Continue using yq v3 for this.

In order to make sure the scripts are using the right version of yq, add
basic validation to them to verify they are running v3 or v4 as required
for their operation.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-30 08:08:08 -07:00
Travis Nielsen d9ac8ce490 test: upgrade integration test from 1.7 to master
The upgrade integration test was from rook v1.6 to the latest master.
This was necessary until we are ready for the v1.8 release, from which
time we want to focus the upgrade testing from v1.7 to the latest
master.

The duplication in the test CRs and other resources is now reduced
by the upgrade calling a thin wrapper to forward a call to the
master version of the resource. When a new feature is added that
needs to be differentiated from the previous version, the method
then can be implemented instead of wrapping the master implementation.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-30 08:07:42 -07:00
Sébastien Han decc481d40 Merge pull request #8740 from leseb/examples-layout
core: change directory layout
2021-11-30 16:02:37 +01:00
Sébastien Han c890710b63 core: change directory layout
As per discussion, proposing a new layout for the charts/yaml/olm files.

./deploy
├── charts
│   ├── rook-ceph
│   │   └── templates
│   └── rook-ceph-cluster
│       └── templates
├── examples
│   ├── csi
│   │   ├── cephfs
│   │   └── rbd
│   ├── flex
│   ├── monitoring
│   ├── pre-k8s-1.16
└── olm
    └── assemble

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-30 09:12:53 +01:00
Olivier Bouffet 5edeff4df8 object: fix search user in objectstore
avoid failed reconcile when multiple multisite objectstore are configured

Signed-off-by: Olivier Bouffet <olivier.bouffet@infomaniak.com>
2021-11-29 20:38:22 +01:00
Sébastien Han 83f7c2be87 Merge pull request #9259 from parth-gr/deviceClass
osd: update existing OSDs with deviceClass
2021-11-29 18:06:08 +01:00
parth-gr cbe505d122 osd: update existing OSDs with deviceClass
If we apply useAllNodes to false for the current deployment,
the OSDs should get updated with the individual nodes values and config,
The deviceClass was not updating to the existing OSDs because there was
bug in the check.
The check osdInfo.DeviceClass == "" which should be
checked like this osdInfo.DeviceClass == "None"

Updated the code so OSDs can make use of the devices present

Signed-off-by: parth-gr <paarora@redhat.com>
2021-11-29 21:40:02 +05:30
Sébastien Han bf662ec7b9 Merge pull request #9264 from leseb/fix-9151
cephfs-mirror: various fixes for random bootstrap peer import errors
2021-11-29 15:45:14 +01:00
Sébastien Han bd962e3e85 Merge pull request #9104 from BlaineEXE/nfs-restart-with-configmap
nfs: restart nfs servers when configmap is updated
2021-11-26 17:06:16 +01:00
Sébastien Han d8a8b05c1f cephfs-mirror: use combined output to possibly catch peer import error
By adding a combined output to the executor we might be able to fetch more
error messages.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-26 15:53:30 +01:00
Sébastien Han 7f7a72d941 core: add the ability to execute ceph commands with a combined output
Sometimes Ceph uses a different standard output to return errors or
merges standard error to standard out. So let's allow some commands to
return both in the output.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-26 15:52:49 +01:00
Sébastien Han d1cdba420a cephfs-mirror: try to mitigate peer import error
Some users have reported issues while adding the token, this is not
always reproducable so perhaps it's a typo when importing the token and
adding trailing spaces.

Closes: https://github.com/rook/rook/issues/9151
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-26 15:45:01 +01:00
Sébastien Han 2519018b26 Merge pull request #9262 from leseb/bot-warn
bot: add bot message when pushing on a release branch
2021-11-26 10:06:33 +01:00
Sébastien Han ae6521f277 Merge pull request #9261 from leseb/master-for-9249
object: fix rgw ceph config
2021-11-26 10:02:59 +01:00
Sébastien Han c2d36c3343 bot: add bot message when pushing on a release branch
Normally, only the mergify bot should send out PR to the release
branches. All contributions should generally go in master first then be
backported. Let's warn to avoid merging PR send against a release
branch.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-25 18:51:38 +01:00
Olivier 561cede1ef object: fix rgw ceph config
use Zone and ZoneGroup instead of storename for rgw_zone and rgw_zonegroup

Signed-off-by: Olivier Bouffet <olivier.bouffet@infomaniak.com>
(cherry picked from commit c92270cd66)
2021-11-25 18:38:42 +01:00
Sébastien Han 0dcfefa3a1 Merge pull request #9258 from leseb/revert-snyk
build: revert "build: images/cross/Dockerfile to reduce vulnerabilities"
2021-11-25 15:59:42 +01:00
Sébastien Han 619f5eee6e build: revert "build: images/cross/Dockerfile to reduce vulnerabilities"
This reverts commit ceec0bc48c. The build
is failing to push images to master. This is not a critical update so
let's remove.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-25 14:53:21 +01:00
Sébastien Han 35b28f760f Merge pull request #9250 from rook/snyk-fix-8a1398b3ebbfb4eeed7c40de45c7b73b
[Snyk] Security upgrade ubuntu from xenial-20210114 to impish-20211015
2021-11-25 13:49:46 +01:00
snyk-bot ceec0bc48c build: images/cross/Dockerfile to reduce vulnerabilities
The following vulnerabilities are fixed with an upgrade:
- https://snyk.io/vuln/SNYK-UBUNTU1604-LIBGCRYPT20-1585790
- https://snyk.io/vuln/SNYK-UBUNTU1604-SYSTEMD-1320131
- https://snyk.io/vuln/SNYK-UBUNTU1604-SYSTEMD-1320131
- https://snyk.io/vuln/SNYK-UBUNTU1604-SYSTEMD-1320131
- https://snyk.io/vuln/SNYK-UBUNTU1604-SYSTEMD-1320131

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-25 11:59:22 +01:00
Sébastien Han 7a223adde4 Merge pull request #9230 from leseb/osd-rm-check
osd: check if osd is safe-to-destroy before removal
2021-11-25 10:55:34 +01:00
Sébastien Han ad2c3c2ae6 Merge pull request #8931 from BlaineEXE/test-rgw-multisite-in-nightly-tests
test: test rgw multisite nightly
2021-11-25 10:25:57 +01:00
Sébastien Han 7402c2cce6 osd: check if osd is ok-to-stop before removal
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.

So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-25 10:21:32 +01:00
Sébastien Han 3e299fbe13 Merge pull request #9236 from leseb/fix-9234
core: fix openshift security context
2021-11-24 18:27:29 +01:00