Commit Graph
7560 Commits
Author SHA1 Message Date
Travis Nielsen 0591753673 docs: update master doc links to latest
With the renaming of master docs to latest, all the links also
need to be updated.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-22 12:41:32 -06:00
Travis Nielsen 0ec1303606 Merge pull request #8729 from Madhu-1/fix-8153
ceph: make provisioner replicas configurable
2021-09-22 11:00:00 -06:00
Sébastien Han e8d540c11f Merge pull request #8785 from leseb/fix-multisite
ci: fix multisite test
2021-09-22 15:37:29 +02:00
Travis Nielsen cd6ee3bce8 Merge pull request #8739 from Madhu-1/reduce-csi-permission
ceph: modify CephFS provisioner permission
2021-09-22 07:31:04 -06:00
Sébastien Han 56b5068712 Merge pull request #8765 from BlaineEXE/fix-object-debug-message
rgw: fix misleading log line in rgw health checker
2021-09-22 11:06:05 +02:00
Sébastien Han 7c8dc4bc1a ci: fix multisite test
We just need to wait a little for the object to be replicated to the
other gateway. A simple retry solves this.

Closes: https://github.com/rook/rook/issues/8671
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-22 10:23:02 +02:00
Madhu Rajanna 95775fd445 ceph: modify CephFS provisioner permission
As like RBD, CephFS provisioner pod need not to
run as privileged. as its not doing any operation
like plugin pods which does mounting and unmounting
removing the permissions for the same.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-09-22 13:30:05 +05:30
Sébastien Han 50fb1b7086 Merge pull request #8756 from humblec/min-version
ceph: lift minimum supported version of ceph csi to v3.3.0
2021-09-22 09:50:40 +02:00
Blaine Gardner c8b26e458c rgw: fix misleading log line in rgw health checker
There was a log line that informed that the object store status would
not be updated because the status was deleting erroneously. Move the
line to the correct position.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-22 07:33:57 +00:00
Madhu Rajanna ed5f281a74 ceph: make provisioner replicas configurable
added new option to set the provisioner replicas.
with this new option the user/admin can choose
how many replicas he want for provisioner pod if
number of nodes is greater than 1.

fixes #8153

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-09-22 11:24:53 +05:30
Humble Chirammal 731b0f5274 ceph: lift minimum supported version of ceph csi to v3.3.0
With the release of Ceph CSI 3.4.0, Ceph CSI project came up
with a new support policy where we support only versions >= 3.3.0

The supported window of Ceph CSI versions is known as "N.(x-1)":
 (N (Latest major release) . (x (Latest minor release) - 1)).

For example, if Ceph CSI latest major version is 3.4.0 today,
support is provided for the versions above 3.3.0.
If users are running an unsupported Ceph CSI version, they will be
asked to upgrade when requesting support for the cluster.

This PR lift the minimum supported version of Ceph CSI to 3.3.0
in this repo.

Fix https://github.com/rook/rook/issues/8709

Ref #
https://github.com/ceph/ceph-csi/releases/tag/v3.4.0
https://github.com/ceph/ceph-csi/#known-to-work-co-platforms

Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
2021-09-22 10:16:43 +05:30
Travis Nielsen 6bcb569a88 Merge pull request #8771 from kubealex/patch-1
helm: Add possibility to default filesystem storageclass in rook-ceph-cluster chart
2021-09-21 13:05:22 -06:00
kubealex d5f42aa4a1 ceph: corrected placement of values
I fixed the placement of the modification, from object storage to filesystem

Signed-off-by: kubealex <al.rossi87@gmail.com>
2021-09-21 20:07:27 +02:00
Travis Nielsen 18707ffdf5 Merge pull request #8772 from leseb/admin-user-secondary-cluster
rgw: do not create the rgw ops user on the secondary cluster
2021-09-21 11:45:17 -06:00
Sébastien Han 8786b40d64 rgw: do not create the rgw ops user on the secondary cluster
If the cluster where the rgw is started is secondary and not primary,
trying to create the admin ops user will fail with:

```
Please run the command on master zone.
Performing this operation on non-master zone
leads to inconsistent metadata between zones
```

So we need to force the creation regardless, it is fine the creation will
return UserAlreadyExist and then we just read the current user.

Closes: https://github.com/rook/rook/issues/8671
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 16:58:03 +02:00
Sébastien Han 470fbfd341 Merge pull request #8743 from leseb/next-pacific
ceph: use next ceph v16.2.6 pacific version
2021-09-21 16:41:45 +02:00
Sébastien Han c1a88f34d4 mds: change init sequence
The MDS core team suggested with deploy the MDS daemon first and then do
the filesystem creation and configuration. Reversing the sequence lets
us avoid spurious FS_DOWN warnings when creating the filesystem.

Closes: #8745
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 15:43:34 +02:00
Alessandro Rossi 4b20c5ce6b Merge branch 'rook:master' into patch-1 2021-09-21 14:49:37 +02:00
kubealex eead60457b ceph: add default field to filesystem-sc helm chart
I Added the chance to default filesystem storageclass in helm chart

Signed-off-by: kubealex <al.rossi87@gmail.com>
2021-09-21 14:49:03 +02:00
Sébastien Han f040c37ad1 ci: fix mirror test with v16.2.6
The new version v16.2.6 has a different behavior when it comes to the
number of cephfs-mirror socket files. Previous version had exactly 3 and
now has like way more...
So let's just check for the presence of more sockets.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 09:16:33 +02:00
Sébastien Han 0c33493f27 ceph: bump manifests to ceph pacific 16.2.6
New version is out so let's use it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 09:16:32 +02:00
Travis Nielsen 404dccdf2d Merge pull request #8721 from jmolmo/8510
ceph: do not use http for mgr liveness probe
2021-09-20 17:05:30 -06:00
Sébastien Han 19837969ea Merge pull request #8722 from leseb/scc-rile
ceph: export security context constraints
2021-09-20 18:43:03 +02:00
Blaine Gardner acbda9381b Merge pull request #8708 from BlaineEXE/health-checker-failure-should-fail-reconcile
ceph: retry object health check if creation fails
2021-09-20 10:27:20 -06:00
Travis Nielsen 8755759a0a Merge pull request #8690 from jmolmo/issue_8669
ceph: fix probable cause of intermittent fails in the manager test
2021-09-20 09:59:54 -06:00
Sébastien Han a8c3b363ee ceph: export security context constraints
Rook now exports both ceph and csi SCCs to run on openshift so any
program pulling rook can inject the right SCCs.
In the meantime, default SCC has been reinforced to be less permissive.
Documentation has been reworked and a new path for openshift
installations has been created. Rook still provides the YAML files for
manual installations.

Closes: #8713
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-20 17:54:48 +02:00
Juan Miguel Olmo Martínez 5c75139581 ceph: fix probable cause of intermittent fails in the manager test
Explicitly set the length of the string parameter in json.Unmarshal method

fixes: https://github.com/rook/rook/issues/8669

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2021-09-20 13:29:45 +00:00
Juan Miguel Olmo Martínez 7fbd9f2225 ceph: do not use http for mgr liveness probe
When private/public network have been defined in the Ceph rook cluster it is not
possible to configure properly the ip address of the liveness probe
for the manager.
Changes in the manager in Pacific introduced this regression.

This change replaces the http probe by a command probe, avoiding thus to
determine what is going to be the ip address of the manager before launching
the pod.

fixes: https://github.com/rook/rook/issues/8510

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2021-09-20 13:28:51 +00:00
Blaine Gardner a357db9bb5 Merge pull request #8736 from BlaineEXE/remove-jenkins-minimal-test-versions
test: remove unused minimal test versions
2021-09-17 17:08:06 -06:00
Blaine Gardner 264256acfb test: remove unused minimal test versions
The `*SuiteMinimalTestVersion` vars in
`tests/integration/ceph_base_deploy_test.go` are no longer needed with
Jenkins no longer being used. Remove them.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 16:49:29 -06:00
Blaine Gardner 5383ba2df2 ceph: retry object health check if creation fails
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 16:24:18 -06:00
Travis Nielsen 5795df728a Merge pull request #8723 from BlaineEXE/separate-object-full-test-suite
test: make object e2e test its own test
2021-09-17 15:57:49 -06:00
Blaine Gardner c60bf241e1 test: make object e2e test its own test
The CephSmokeSuite is becoming quite large and long, and most of the
length is now related to the object e2e tests. Separate the full object
e2e test into CephObjectSuite, and only test the object 'lite' test in
the CephSmokeSuite.

Resolves #8714

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 15:27:27 -06:00
Travis Nielsen 0d8fd9d8a4 Merge pull request #8607 from leseb/refact
ceph: refactor operator initialization sequence and add more controllers (csi, lib-bucket-prov, discover daemon, flex)
2021-09-17 11:38:35 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 78ef8768c2 Merge pull request #8738 from leseb/yaml-parse
ceph: remove unnecessary package
2021-09-17 16:48:09 +02:00
Sébastien Han 68a4bc2d1b ceph: remove unnecessary package
We don't need to use github.com/ghodss/yaml since
"k8s.io/apimachinery/pkg/util/yaml" provides the same functionality and
we already import it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:10:47 +02:00
Sébastien Han 32984f7297 Merge pull request #8746 from leseb/force-16.2.5
ci: force a particular ceph version
2021-09-17 16:09:12 +02:00
Travis Nielsen b3f5ec6dcf Merge pull request #8732 from leseb/cephfs-mirror-doc
doc: fix cephfs-mirror documentation
2021-09-17 07:54:41 -06:00
Sébastien Han ae291afb2f ci: force a particular ceph version
Let's force v16.2.5 since the CI is broken with 16.2.6. This gives us
time to continue to merge work and work on fixing deployments with
16.2.6 in parallel.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 15:28:48 +02:00
Sébastien Han fd28497cb4 docs: fix cephfs-mirror documentation
The steps to configure the peers are detailed in the CephFilesystem
section. Only the CephFilesystem is holding the peer configuration, not
the CephFilesystemMirror which only controls the bootstrap of the
daemon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 09:50:13 +02:00
Travis Nielsen e2e3fc83fc Merge pull request #8673 from leseb/stale-bot
bot: change issue and pr stale days
2021-09-16 08:31:31 -06:00
Blaine Gardner 79e8cf0eea Merge pull request #8715 from BlaineEXE/prune-unused-crd-permission-in-clusterrole
ceph: remove unused crd permission from roles
2021-09-16 08:04:58 -06:00
Sébastien Han 875aec74cc bot: bump action to v4
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-16 10:04:56 +02:00
Sébastien Han 796abd7805 bot: change issue and pr stale days
Let's be a bit more aggressive to ensure a consistent amount of
in-progress PRs and issues.
So now PR will go stale after 30 days and issues after 60 days.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-16 10:04:46 +02:00
Sébastien Han bcae365cb9 Merge pull request #8698 from sp98/fix-pdb-reconcile
ceph: reconcile osd pdb if allowed disruption is 0
2021-09-15 10:47:54 +02:00
Sébastien Han 1632bfaa5a Merge pull request #8711 from subhamkrai/dry-run-check
ci: dry-run is deprecated and replaced with --dry-run=client
2021-09-15 10:06:48 +02:00
Sébastien Han c4a6473fc8 Merge pull request #8435 from BlaineEXE/upgrade-doc-add-mirroring-changes
docs: ceph: add peer spec migration to upgrade doc
2021-09-15 09:26:12 +02:00
Travis Nielsen 7f3eebe578 Merge pull request #8695 from travisn/master-docs-latest
docs: Publish master docs to latest path
2021-09-14 14:27:56 -06:00
Blaine Gardner 7c511328e4 docs: update rbd mirror docs for block pool config
Remove legacy documentation for configuring RBD mirroring. While we
still support legacy mirroring configs, we want to encourage new users
to use the CephBlockPool configuration for mirroring.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-14 13:39:28 -06:00