Commit Graph
76 Commits
Author SHA1 Message Date
Sébastien Han ebbf332d8d ceph: update rgw and mds deployment for logCollector
If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.

Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-22 18:19:07 +01:00
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
Julien Girardin d63d9b7d4d ceph: change external rgw detection, not relying on cluster
After #6217, internal rgw spawning on external cluster was still broken,
operator claiming:
```
ceph-object-controller: failed to reconcile failed to create object
store deployments: failed to reconcile external endpoint:
failed to create or update object store "arch-cloud" endpoint:
failed to create endpoint "rook-ceph-rgw-arch-cloud". Endpoints
"rook-ceph-rgw-arch-cloud" is invalid: subsets[0]:
Required value: must specify `addresses` or `notReadyAddresses`
```

To spawn internal rgw for external cluster, changes detection of
what mean 'internal' for rgw and only use the externalRgwEndpoints
list for that independently of the status of the cluster
Now there is multiple posibilities:

 * internal cluster and internal rgw pods: "normal case"
 * external cluster and internal rgw pods: <= now working with the PR
 * external cluster and external rgw pods: the external case
 * internal cluster and external rgw pods: <= new case that could exist

Signed-off-by: Julien Girardin <jugirardin@free.fr>
2021-01-08 11:28:41 +01:00
Blaine Gardner b8dc2a3214 ceph: update lib bucket provisioner
Update to the latest lib bucket provisioner code.
Fixes issue 6650

Modifies CRD for objectbucketclaims to fix an additional bug where an
ObjectBucket's 'ClaimRef' is lost due to the CRD validation being
specified incorrectly.

Changes OBC deletion/cleanup to delete the bucket before the user. A
user cannot be deleted without an unsafe purge option if the user has
buckets associated to it.

Does not reintroduce bug 6767 from previous fix for 6650

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-18 12:48:54 -07:00
Blaine Gardner 22cd1bf0f2 Revert "ceph: update object bucket provisioner library"
This reverts commit 3b4ee6c8e3.

Revert the fix that introduces the following bug on upgrade:
https://github.com/rook/rook/issues/6767

This fix as well as a fix for the upgrade case will follow in days to
come.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-04 10:04:34 -07:00
Blaine Gardner 3b4ee6c8e3 ceph: update object bucket provisioner library
The library for object bucket provisioning is updated to fix errors
during provisioning.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-03 09:43:12 -07:00
Sébastien Han 97be23e374 ceph: apply finalizer before updating object status
We must add the finalizer right after the object creation otherwise the
seerver will later return an error on update that the object has been
modified. Indeed, it has been by the task that updates the status when
the object is first created.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-18 18:16:59 +01:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Blaine Gardner 98d5c73d5e ceph: fill in rgw deployment version
The RGW deployment's ceph-version label should now display the same
version as the image it is using.

It previously displayed an empty version: 0.0.0-0.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-11-11 15:13:31 -07:00
Sébastien Han e3d032dbc1 Merge pull request #6489 from travisn/stretch-mons
ceph: Stretch cluster configuration
2020-11-06 15:16:39 +01:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
Travis Nielsen e72b69f8e0 ceph: validate pools for stretch clusters
Pools in stretch clusters will all use the same crush rule that
is generated by the operator when configuring the cluster for
stretch mode, with replica 4 and two replicas per failure domain.
Pools cannot create new crush rules, so we require that replica be
4 when the pools is created, EC is not allowed, and no new
crush rule will be created.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-05 09:23:57 -07:00
Jonas Schäfer 7ace6ac255 ceph: make root= CRUSH label value configurable
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.

By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.

To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.

The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.

Fixes #4993.

Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>
2020-10-29 09:22:41 +01:00
Ali Maredia 8964941d44 ceph: add unit tests for object store controller with multisite
Add unit tests for when the "Zone" is configured
in an object store so that the object store joins
the CephObjectZone and the zone's corresponding
multisite configuration.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-10-22 11:36:06 -04:00
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
subhamkrai 34d27d55d5 ceph: update cephobjectstore monitoring logic
currently, cephobjectstore monitoring is never
stopped for the following flow:
 1. objectstore is deleted
 2. go routine is stopped
 3. objectstore is created
 4. objectstore is deleted
so, removing clusterResourceDeleted flag and
working with the map only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-03 15:58:48 +05:30
subhamkrai e21854145a ceph: disable object store port when secureport specified
currently, the object store schema requires the
port to be non-zero. This does not allow http
access to the object store to be disabled.
so setting the port minimum to 0 to from 1 in the cr.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-08-12 21:57:01 +05:30
subhamkrai acc4ed5df9 ceph: handling gosec errors that are not checked
this commit handles all the gosec g104
i.e audit errors not checked inside pkg.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-24 20:02:06 +05:30
Sébastien Han 8729a206f3 ceph: let the healthcheck set the status phase
Instead of setting the Phase of the CR once the reconcile is done, let's
actually set in from the healthcheck so it is more accurate.

Closes: https://github.com/rook/rook/issues/5249
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-22 17:25:12 +02:00
Sébastien Han 8f4e37b80a Merge pull request #5452 from leseb/osd-wipe
ceph: enhancement disk sanitize options
2020-07-22 16:11:35 +02:00
Ali Maredia e6ed4ff8ea ceph: minor fixes + add realm/zg/zone to object context
- Add realm/zone group/zone to Object Context,
so that any call to runAdminCommand has the
same realm, zone group, and zone as the object
store that the call is going to.

- Also modify the delete code for the object-store
using multisite to remove the endpoints of an
object store from a zone instead of the normal
clean-up

- Added more debug logging all around the object-store
code related to multisite

- files generated by rerun of `make codegen`

- change back the edit on the Copyright in object.go

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-07-21 16:47:59 -04:00
Ali Maredia 0f7e55780c ceph: add pool creation for ceph-object-zone
When zones are created the ceph RGWs inside the
those zones should be using pools with the zones
name not the object-store's name

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-07-21 13:30:41 -04:00
Ali Maredia 5d9c2f09a9 ceph: enable the realm pull on ceph-object-realm
This commit:
- adds the pull section on the CephObjectRealm spec
to enables realms to be pulled instead of created.

- creates the system user and generates the access
key and secret key for the user and zones in a realm.

- adds "omitempty" to fields in multisite CRs that
are not required

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-07-21 13:30:19 -04:00
Sébastien Han 0a6cce43fb ceph: add clusterResourceDeleted flag
During the deletion sequence of the CR, we now set a success flag at the
end, right before removing the finalizer.
This avoids the case where we fail to remove the finalizer.

2020-07-21 13:29:46.334816 E | ceph-object-controller: failed to reconcile failed to remove finalizer: failed to remove finalizer "cephobjectstore.ceph.rook.io" on "teststore": Operation cannot be fulfilled on cephobjectstores.ceph.rook.io "teststore": the object has been modified; please apply your changes to the latest version and try again

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-21 17:11:29 +02:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Travis Nielsen 4681f9e73d ceph: consolidate ceph config and client packages
The ceph config and client packages are conceptually the same.
To avoid circular dependencies in some cases, we simplify by combining
the packages into the client package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:43 -06:00
Sébastien Han ccb52b84e6 ceph: configurable status checks and livenessprobe
This commit allows us to configure status check for each daemon:

* "mon": health check on the ceph monitors (quorum)
* "osd": health check on the ceph osds
* "status": ceph health status check

Each check is controlled by the following settings:

* disabled: whether to disable the check (default: false)
* internal: interval to run the check
* timeout: only valid for mons, is the timeout for unresponsive mon
before failling over.

Example to disable the status health check:

```yaml
healthCheck:
  daemonHealth:
    status:
      disabled: true
```

As part of that, pod's livenessprobe can now be configured via the
following settings:

* disabled: whether to enable or not
* probe: override the current probe in place by a new one

```yaml
healthCheck:
  livenessProbe:
    mon:
      disabled: true
```

Closes: https://github.com/rook/rook/issues/5772
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-10 19:21:59 +02:00
Sébastien Han d097343001 Merge pull request #5774 from leseb/object-user-endpoint
ceph: add s3 endpoint to the rgw user cr
2020-07-10 17:22:22 +02:00
Sébastien Han 5c5008e27c ceph: add s3 endpoint to CephObjectStore CRD
For the CephObjectStore, the Status field of the CR has a new property
called "info", which contains a couple of information about the CR.
The first new addition is the endpoint of the object store.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-10 12:06:25 +02:00
Jiffin Tony Thottan 2308ad311e ceph: check for ceph object user cleanup
Do not delete ceph object store until ceph object users are cleaned up

Closes: #5712
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-07-08 23:46:30 +05:30
Sébastien Han e4eaa91ede ceph: add rgw endpoint healthcheck
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.

A good status will look like:

status:
  endpointStatus:
    lastChanged: "2020-06-25T13:47:45Z"
    lastChecked: "2020-06-25T13:48:46Z"
  phase: Connected

A failed status:

status:
  endpointStatus:
    details: |-
      error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
      caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
    health: ERROR

This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.

Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-02 16:34:10 +02:00
Sébastien Han c7f255a0e8 ceph: add the ability to set any pool property
We can now explicitly set any property on a given pool by using the new
Property field in the CephBlockPool Spec.

Also, this fixes the case where both `CephBlockPool` and `CephCluster`
are created at the same time. When Rook creates the pool, the cluster is
still being bootstrapped and the global option
`osd_pool_default_pg_autoscale_mode` has not bee set yet. So the pool
gets created but its `pg_autoscale_mode` property is set to `warn`
instead of `on`.

Closes: https://github.com/rook/rook/issues/5608V
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-17 10:08:00 +02:00
Ali Maredia fc579f4520 ceph: initial commit for ceph rgw multisite resources
This commit contains CR implementations for:
CephObjectRealm
CephObjectZoneGroup
CephObjectZone

Also there are changes made to the objectstore
to add rgws in the object-store to zones and
zone groups in a multisite configuration and
the removal of the --default parameter for any
realms/zonegroups/zones that are created.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-06-09 16:28:53 -04:00
Santosh PillaiandElise Gafford 5ef88f8f09 ceph: finalizer for OBC cleanup
This PR is for rebasing #4683 with master. With new controller runtime changes most the changes with #4683 got redundant.
Only valid change is waiting for all the object buckets to be cleaned up before removing the finalizer.

Co-authored-by: Elise Gafford <egafford@redhat.com>
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-05-18 23:54:59 +05:30
Madhu Rajanna 81688398f2 cleanup: use err.Wrap when the formatting is not required
Replaced err.Wrapf with err.Wrap when the formatting
is not required.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-04-29 17:42:59 +05:30
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Travis Nielsen 269cfe1199 ceph: update status on current version of resources
When updating the status on resources sometimes the update fails due to
the resource being an outdated version. This frequently occurs when
the same reconcile loop updates the status multiple times, or the finalizer
is added, or some other update to the resource. Upon the next reconcile
the error would generally go away since the status or finalizer didn't
need to be updated multiple times. But now the error is not expected
since the controller will retrieve the latest version of the resource
immediately before attempting to update it.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-09 11:20:45 -06:00
Sébastien Han 040193bb5a ceph: remove DaemonType type
This type was a string already and was just making us doing string()
calls all the time to it's not worth it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-01 09:08:18 +02:00
Sébastien Han f136105951 ceph: use a different port for rgw on sdn
When running on SDN, the container is not privileged and the network
stack is not exposed so running the gw on port 80 will not work (not
enough privileged).

Closes: https://github.com/rook/rook/issues/5106
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-30 19:37:11 +02:00
Travis Nielsen 5254de1a8c ceph: simplify pool model to v1 types
The pools had some legacy structs that translated between the
ceph.v1 types used by the CRDs and the internal implementation
of the pools. This simplifies the pool implementation by removing
the intermediate model and leaving us only with the ceph v1
pool types and no unnecessary translation.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-25 16:59:12 -06:00
Sébastien Han 8db14885b5 ceph: controller fix misleading debug log
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 19:12:17 +01:00
Sébastien Han a699512911 ceph: rollback to objectmatcher 1.1.0
With 1.1.1, the object matcher seems not to preserve annotation's order,
this is tracked here: https://github.com/banzaicloud/k8s-objectmatcher/issues/24

1.1.0 does not have that issue, so let's go back with 1.1.0.

Closes: https://github.com/rook/rook/issues/5047
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Travis Nielsen 5d2db6a9f1 ceph: allow creation of object store with pre-existing pools
The pools may already be created by the mgr module. If the pool specs
are not specified in the object CR, skip pool creation for the
object store. The pools are required to exist and the reconcile
will fail until they do exist.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-23 23:54:20 -06:00
Sébastien Han f64500a954 ceph: failed reconcile if ceph version is not found
This is critical to get this information so that the label can be set
properly and also the configuration done appropriately.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-20 17:38:03 +01:00
Sébastien Han d3d4cd2f20 ceph: child controller-runtime stop watching for deployment
The bump from 0.2 to 0.4 controller-runtime versions which happened
during the use of Go Modules (bumping Go version to 1.13) introduced an
issue where the deployment annotations keep getting changed during
watch update event.
The easy fix is to stop watching Deployment objects for now.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-19 15:19:23 +01:00
Sébastien Han f268c897e9 ceph: Convert the Ceph ObjectStore controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Sébastien Han a3068dee0b ceph: separate controller for CephBlockPool CRD
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-02 17:31:03 +01:00
Nizamudeen 53883f68cf ceph: Handling Unhandled errors
This commit is to handle all those unhandled errors which raises the gosec warning.

Fixed G104: Unhandled Errors are handled now

Signed-off-by: Nizamudeen <nia@redhat.com>
2020-02-21 22:48:24 +05:30
Travis Nielsen 41fe8c3a2b ceph: call the ParentClusterChanged in a goroutine
When the cluster is upgraded, the child controllers are notified with the
callback to the ParentClusterChanged() method. These should be made
in a separate goroutine so that the cluster controller is not blocked
on the child controllers and also so child controllers aren't blocked
on each other for the upgrade.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-17 16:24:28 -07:00
Sébastien Han 92c1696bfc ceph: cleanup/trim isUpgrade variable a bit
We now always perform ok-to-stop change when the deployment is going to
be updated. This could be for various reasons like change in the spec or
a new Ceph image.
In any case, if Ceph resources are going to be restarted this must be
done carefully.

Also trimmed the isUpgrade variable, the upcoming commit will remove the
need of that variable.

Keeping isUpgrade from the controller code since we need it to detect
whether an orchestration upgrade should be triggered or not (based on
the ceph cluster status).

Also, we keep it on the child controllers so we don't go into an
orchestration phase if the cluster image hasn't changed. The child does
not need to be updated/go into its orchestration phase because the
cluster CR changed.

Closes: https://github.com/rook/rook/issues/4379
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-24 17:02:37 +01:00