Commit Graph
2376 Commits
Author SHA1 Message Date
Travis Nielsen 0ec1303606 Merge pull request #8729 from Madhu-1/fix-8153
ceph: make provisioner replicas configurable
2021-09-22 11:00:00 -06:00
Travis Nielsen cd6ee3bce8 Merge pull request #8739 from Madhu-1/reduce-csi-permission
ceph: modify CephFS provisioner permission
2021-09-22 07:31:04 -06:00
Sébastien Han 56b5068712 Merge pull request #8765 from BlaineEXE/fix-object-debug-message
rgw: fix misleading log line in rgw health checker
2021-09-22 11:06:05 +02:00
Madhu Rajanna 95775fd445 ceph: modify CephFS provisioner permission
As like RBD, CephFS provisioner pod need not to
run as privileged. as its not doing any operation
like plugin pods which does mounting and unmounting
removing the permissions for the same.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-09-22 13:30:05 +05:30
Sébastien Han 50fb1b7086 Merge pull request #8756 from humblec/min-version
ceph: lift minimum supported version of ceph csi to v3.3.0
2021-09-22 09:50:40 +02:00
Blaine Gardner c8b26e458c rgw: fix misleading log line in rgw health checker
There was a log line that informed that the object store status would
not be updated because the status was deleting erroneously. Move the
line to the correct position.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-22 07:33:57 +00:00
Madhu Rajanna ed5f281a74 ceph: make provisioner replicas configurable
added new option to set the provisioner replicas.
with this new option the user/admin can choose
how many replicas he want for provisioner pod if
number of nodes is greater than 1.

fixes #8153

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2021-09-22 11:24:53 +05:30
Humble Chirammal 731b0f5274 ceph: lift minimum supported version of ceph csi to v3.3.0
With the release of Ceph CSI 3.4.0, Ceph CSI project came up
with a new support policy where we support only versions >= 3.3.0

The supported window of Ceph CSI versions is known as "N.(x-1)":
 (N (Latest major release) . (x (Latest minor release) - 1)).

For example, if Ceph CSI latest major version is 3.4.0 today,
support is provided for the versions above 3.3.0.
If users are running an unsupported Ceph CSI version, they will be
asked to upgrade when requesting support for the cluster.

This PR lift the minimum supported version of Ceph CSI to 3.3.0
in this repo.

Fix https://github.com/rook/rook/issues/8709

Ref #
https://github.com/ceph/ceph-csi/releases/tag/v3.4.0
https://github.com/ceph/ceph-csi/#known-to-work-co-platforms

Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
2021-09-22 10:16:43 +05:30
Sébastien Han 8786b40d64 rgw: do not create the rgw ops user on the secondary cluster
If the cluster where the rgw is started is secondary and not primary,
trying to create the admin ops user will fail with:

```
Please run the command on master zone.
Performing this operation on non-master zone
leads to inconsistent metadata between zones
```

So we need to force the creation regardless, it is fine the creation will
return UserAlreadyExist and then we just read the current user.

Closes: https://github.com/rook/rook/issues/8671
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 16:58:03 +02:00
Sébastien Han 470fbfd341 Merge pull request #8743 from leseb/next-pacific
ceph: use next ceph v16.2.6 pacific version
2021-09-21 16:41:45 +02:00
Sébastien Han c1a88f34d4 mds: change init sequence
The MDS core team suggested with deploy the MDS daemon first and then do
the filesystem creation and configuration. Reversing the sequence lets
us avoid spurious FS_DOWN warnings when creating the filesystem.

Closes: #8745
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 15:43:34 +02:00
Sébastien Han 0c33493f27 ceph: bump manifests to ceph pacific 16.2.6
New version is out so let's use it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-21 09:16:32 +02:00
Travis Nielsen 404dccdf2d Merge pull request #8721 from jmolmo/8510
ceph: do not use http for mgr liveness probe
2021-09-20 17:05:30 -06:00
Juan Miguel Olmo Martínez 7fbd9f2225 ceph: do not use http for mgr liveness probe
When private/public network have been defined in the Ceph rook cluster it is not
possible to configure properly the ip address of the liveness probe
for the manager.
Changes in the manager in Pacific introduced this regression.

This change replaces the http probe by a command probe, avoiding thus to
determine what is going to be the ip address of the manager before launching
the pod.

fixes: https://github.com/rook/rook/issues/8510

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2021-09-20 13:28:51 +00:00
Blaine Gardner 5383ba2df2 ceph: retry object health check if creation fails
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-09-17 16:24:18 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 68a4bc2d1b ceph: remove unnecessary package
We don't need to use github.com/ghodss/yaml since
"k8s.io/apimachinery/pkg/util/yaml" provides the same functionality and
we already import it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:10:47 +02:00
Sébastien Han bcae365cb9 Merge pull request #8698 from sp98/fix-pdb-reconcile
ceph: reconcile osd pdb if allowed disruption is 0
2021-09-15 10:47:54 +02:00
Santosh Pillai 7480f6ba62 ceph: reconcile osd pdb if allowed disruption is 0
Rook checks for down OSDs by checking the `ReadyReplicas` count
in the OSD deployement. When an OSD pod goes into CBLO due to
disk failure, there is a delay before this `ReadyReplicas` count
becomes 0. The deplay is very small but may result in rook missing
OSD down event. As a result no blocking PDBs will be created and
only default PDB with `AllowedDisruptions` count as 0 is available.
This PR tries to solve this. The OSD pdb reconciler will be
reconciled again if `AllowedDisruptions` count in the main
PDB is 0.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-09-14 15:56:59 +05:30
Blaine Gardner eed91efdc2 Merge pull request #8667 from thotz/enhanceusernits
ceph: addressing nits from #8211
2021-09-09 12:10:51 -06:00
Sébastien Han 031be4be19 Merge pull request #8675 from subhamkrai/ok-continue
ceph: modify the log info when ok to continue fails
2021-09-09 17:22:13 +02:00
subhamkraiandZeaone 0657804491 ceph: modify the log info when ok to continue fails
correct typo in logging, it was showing `ok-to-stop`
instead of `ok-to-continue` when 'continueUpgradeAfterChecksEvenIfNotHealthy' is true

Co-Authored-by: Zeaone <zeaone@ZeaonedeMacBook-Pro.local>
Signed-off-by: subhamkrai <srai@redhat.com>
2021-09-09 19:42:06 +05:30
Jiffin Tony Thottan 50ecff8f13 ceph: addressing nits from #8211
Addressing remaining nits from the PR #8211

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-09 12:57:32 +05:30
Blaine Gardner 5b0785dde0 Merge pull request #8211 from thotz/enhancecephuser
ceph: add options for cephobjectstore user
2021-09-08 11:57:25 -06:00
Travis Nielsen af8b4c116b Merge pull request #8653 from JrCs/externalip-failover
ceph: use node externalIP if no internalIP defined
2021-09-08 09:07:25 -06:00
JrCs c606f4c488 ceph: use node externalIP if no internalIP defined
In some cases node internalIP is not defined. Then use externalIP if it
exists.

Signed-off-by: JrCs <90z7oey02@sneakemail.com>
2021-09-08 16:26:05 +02:00
Sébastien Han 4cb9b53487 ceph: fix pool deletion when running on multus
The ceph-block-pool controller was missing the network spec from the
cephcluster that contains all the details about networking.
So when the pool was deleting on multus the controller was not proxying
the rbd command to the mgr pod but executed the command in the operator
pod which does not have the network annotation and then connect to the
ceph cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-08 16:10:31 +02:00
Jiffin Tony Thottan ca43800119 ceph: add options for cephobjectstore user
Adding options for quota, bucket limit, caps for the
`cephobjectstoreuser`.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-07 22:43:09 +05:30
Sébastien Han 507c51baaa Merge pull request #8624 from leseb/debug-object
ci: fix object store test
2021-09-07 17:57:40 +02:00
Sébastien Han 8f42bee563 ci: fix object store test
We just need to wait longer when the status is not ready. We needed
another sleep otherwise the status was never nil and the loop went too
fast. See:

```
2021-09-02 16:29:11.372249 I | integrationTest:
2021-09-02 16:29:11.374427 I | integrationTest:
2021-09-02 16:29:11.377764 I | integrationTest:
2021-09-02 16:29:11.379950 I | integrationTest:
2021-09-02 16:29:11.382084 I | integrationTest:
2021-09-02 16:29:11.385383 I | integrationTest:
2021-09-02 16:29:11.388499 I | integrationTest:
2021-09-02 16:29:11.391301 I | integrationTest:
2021-09-02 16:29:11.393545 I | integrationTest:
2021-09-02 16:29:11.396249 I | integrationTest:
```

Signed-off-by: Sébastien Han <seb@redhat.com>

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-07 17:15:20 +02:00
Travis Nielsen 1cb97574ea ceph: allow an even number of mons
While an even number of mons can cause lower availability of
mon quorum, it also can provide higher durability for the cluster.
Mon quorum can be restored from a single mon according to the
disaster recovery guide, so there may be scenarios where
an even number of mons may be preferable.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-03 07:43:32 -06:00
parth-gr e50d83771e ceph: add pdb for mgr
Currently PDB for MGR is not been created,
Creating the PDB for MGR if its count is 2

Closes: https://github.com/rook/rook/issues/8275
Signed-off-by: parth-gr <paarora@redhat.com>
2021-09-03 14:30:58 +05:30
Travis Nielsen 8af24cf6d1 Merge pull request #8629 from llamerada-jp/ceph-fix-error-message
ceph: fix error message in UpdateNodeStatus
2021-09-02 11:17:10 -06:00
Sébastien Han 072c8f43df ceph: do not reconcile if the op is not ready
When the op is initializing, the ceph config is not ready and thus we
should reconcile. This avoids a double reconcile.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-02 13:50:32 +02:00
Hiroya Onoe 2ff5413b75 ceph: fix error message in UpdateNodeStatus
The error message in UpdateNodeStatus regards the second argument
as node name. However, it is a PVC name in OSD on PVC.

Signed-off-by: Hiroya Onoe <onoehiroya@gmail.com>
2021-09-02 02:43:16 +00:00
Travis Nielsen afd809e894 Merge pull request #8615 from llamerada-jp/ceph-fix-set-owner-references
ceph: avoid duplicate ownerReferences
2021-09-01 07:16:08 -06:00
Blaine Gardner a1814af1d9 ceph: remove NFS and Cassandra operator code
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-08-31 14:07:02 -06:00
Sébastien Han d4dd0577f9 Merge pull request #8529 from sp98/id-mapping
ceph: add ClusterID and PoolID mappings between local and peer cluster
2021-08-31 17:32:14 +02:00
Santosh Pillai 3f8abec403 ceph: add ClusterID and PoolID mappings between local and peer cluster
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters

This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-08-31 20:08:56 +05:30
Sébastien Han d675969567 ceph: fix vault kv secret engine auto-detection
Passing struct by value essentially gives you a copy, so when modified
within a function, the scope is then reduced to that function. Using
pointers solves that you mutate the struct as many times as you want from
anywhere.
As a result, the auto-detection of the Vault KV backend was not working
correctly.
Also, added a ton of unit tests for Vault.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-31 15:29:04 +02:00
YZ775andHiroya Onoe c38e1f1ddd ceph: avoid duplicate ownerReferences
If we call OwnerInfo.SetOwnerReference for an object multiple times,
it results in OwnerReference duplication.

Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Co-authored-by: Hiroya Onoe <onoehiroya@gmail.com>
2021-08-31 04:58:59 +00:00
Travis Nielsen 7cfae42a62 ceph: set the filesystem status when mirroring not enabled
When mirroring is enabled on the filesystem, the status was not
being set on the filesystem. Now the reconcile will ensure the
status is updated on the CR whether or not mirroring is enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-08-27 17:28:25 -06:00
Travis Nielsen 42717d2fca Merge pull request #8566 from subhamkrai/merge-tolerations
ceph: merge toleration for osd/prepareOSD pod if specified both places
2021-08-26 08:55:27 -06:00
Travis Nielsen b528310b79 Merge pull request #8590 from travisn/mon-store-config
ceph: Consolidate the calls to set mon config
2021-08-25 08:31:20 -06:00
Travis Nielsen edb8aa1834 Merge pull request #8584 from parth-gr/string-instance1
core: convert util.NewSet() to sets.NewString()
2021-08-25 08:30:24 -06:00
subhamkrai e1f232eeb7 ceph: merge toleration for osd/prepareOSD pod if specified both places
earlier, `ApplyToPodSpec()` was only taking one toleration and ignoring
tolerations from `placement.ALL()`.

this commit merge toleration for Mgr,Mon,Osd pod
example, for osd it will merge
spec.placement.all and
storageDeviceClassSets.Placement(in case of pvc) or
spec.placement.osd(in case of non-pvc's).

Signed-off-by: subhamkrai <srai@redhat.com>
2021-08-25 19:34:28 +05:30
parth-gr 77ba1ef11b core: convert util.NewSet() to sets.NewString()
Converting util.NewSet() instance to use sets.NewString() instance

Closes: https://github.com/rook/rook/issues/8479
Signed-off-by: parth-gr <paarora@redhat.com>
2021-08-25 13:13:05 +05:30
Sébastien Han 41e915d411 Merge pull request #8493 from leseb/admission-controller
ceph: move the admission webhook to the operator
2021-08-25 09:30:54 +02:00
Travis Nielsen cdfe1982d1 ceph: consolidate the calls to set mon config
The mon config had two different implementations that have evolved
over the lifetime of the project. This is a simple refactor to remove
the SetConfig() option and stick with the MonStore as a single
implementation for updating the mon store.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-08-24 14:38:26 -06:00
Travis Nielsen 7af10af781 Merge pull request #8514 from thotz/obcupdateapi
ceph: add support for update() from lib-bucket-provisioner
2021-08-24 13:28:46 -06:00