If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.
Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
After #6217, internal rgw spawning on external cluster was still broken,
operator claiming:
```
ceph-object-controller: failed to reconcile failed to create object
store deployments: failed to reconcile external endpoint:
failed to create or update object store "arch-cloud" endpoint:
failed to create endpoint "rook-ceph-rgw-arch-cloud". Endpoints
"rook-ceph-rgw-arch-cloud" is invalid: subsets[0]:
Required value: must specify `addresses` or `notReadyAddresses`
```
To spawn internal rgw for external cluster, changes detection of
what mean 'internal' for rgw and only use the externalRgwEndpoints
list for that independently of the status of the cluster
Now there is multiple posibilities:
* internal cluster and internal rgw pods: "normal case"
* external cluster and internal rgw pods: <= now working with the PR
* external cluster and external rgw pods: the external case
* internal cluster and external rgw pods: <= new case that could exist
Signed-off-by: Julien Girardin <jugirardin@free.fr>
Update to the latest lib bucket provisioner code.
Fixes issue 6650
Modifies CRD for objectbucketclaims to fix an additional bug where an
ObjectBucket's 'ClaimRef' is lost due to the CRD validation being
specified incorrectly.
Changes OBC deletion/cleanup to delete the bucket before the user. A
user cannot be deleted without an unsafe purge option if the user has
buckets associated to it.
Does not reintroduce bug 6767 from previous fix for 6650
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
We must add the finalizer right after the object creation otherwise the
seerver will later return an error on update that the object has been
modified. Indeed, it has been by the task that updates the status when
the object is first created.
Signed-off-by: Sébastien Han <seb@redhat.com>
The RGW deployment's ceph-version label should now display the same
version as the image it is using.
It previously displayed an empty version: 0.0.0-0.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
Pools in stretch clusters will all use the same crush rule that
is generated by the operator when configuring the cluster for
stretch mode, with replica 4 and two replicas per failure domain.
Pools cannot create new crush rules, so we require that replica be
4 when the pools is created, EC is not allowed, and no new
crush rule will be created.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.
By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.
To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.
The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.
Fixes#4993.
Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>
Add unit tests for when the "Zone" is configured
in an object store so that the object store joins
the CephObjectZone and the zone's corresponding
multisite configuration.
Signed-off-by: Ali Maredia <amaredia@redhat.com>
this commit will enable one more linter ineffassign
in golangci-lint.
This linter throws an error when variable is assigned and never used.
`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
currently, cephobjectstore monitoring is never
stopped for the following flow:
1. objectstore is deleted
2. go routine is stopped
3. objectstore is created
4. objectstore is deleted
so, removing clusterResourceDeleted flag and
working with the map only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
currently, the object store schema requires the
port to be non-zero. This does not allow http
access to the object store to be disabled.
so setting the port minimum to 0 to from 1 in the cr.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
Instead of setting the Phase of the CR once the reconcile is done, let's
actually set in from the healthcheck so it is more accurate.
Closes: https://github.com/rook/rook/issues/5249
Signed-off-by: Sébastien Han <seb@redhat.com>
- Add realm/zone group/zone to Object Context,
so that any call to runAdminCommand has the
same realm, zone group, and zone as the object
store that the call is going to.
- Also modify the delete code for the object-store
using multisite to remove the endpoints of an
object store from a zone instead of the normal
clean-up
- Added more debug logging all around the object-store
code related to multisite
- files generated by rerun of `make codegen`
- change back the edit on the Copyright in object.go
Signed-off-by: Ali Maredia <amaredia@redhat.com>
When zones are created the ceph RGWs inside the
those zones should be using pools with the zones
name not the object-store's name
Signed-off-by: Ali Maredia <amaredia@redhat.com>
This commit:
- adds the pull section on the CephObjectRealm spec
to enables realms to be pulled instead of created.
- creates the system user and generates the access
key and secret key for the user and zones in a realm.
- adds "omitempty" to fields in multisite CRs that
are not required
Signed-off-by: Ali Maredia <amaredia@redhat.com>
During the deletion sequence of the CR, we now set a success flag at the
end, right before removing the finalizer.
This avoids the case where we fail to remove the finalizer.
2020-07-21 13:29:46.334816 E | ceph-object-controller: failed to reconcile failed to remove finalizer: failed to remove finalizer "cephobjectstore.ceph.rook.io" on "teststore": Operation cannot be fulfilled on cephobjectstores.ceph.rook.io "teststore": the object has been modified; please apply your changes to the latest version and try again
Signed-off-by: Sébastien Han <seb@redhat.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ceph config and client packages are conceptually the same.
To avoid circular dependencies in some cases, we simplify by combining
the packages into the client package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit allows us to configure status check for each daemon:
* "mon": health check on the ceph monitors (quorum)
* "osd": health check on the ceph osds
* "status": ceph health status check
Each check is controlled by the following settings:
* disabled: whether to disable the check (default: false)
* internal: interval to run the check
* timeout: only valid for mons, is the timeout for unresponsive mon
before failling over.
Example to disable the status health check:
```yaml
healthCheck:
daemonHealth:
status:
disabled: true
```
As part of that, pod's livenessprobe can now be configured via the
following settings:
* disabled: whether to enable or not
* probe: override the current probe in place by a new one
```yaml
healthCheck:
livenessProbe:
mon:
disabled: true
```
Closes: https://github.com/rook/rook/issues/5772
Signed-off-by: Sébastien Han <seb@redhat.com>
For the CephObjectStore, the Status field of the CR has a new property
called "info", which contains a couple of information about the CR.
The first new addition is the endpoint of the object store.
Signed-off-by: Sébastien Han <seb@redhat.com>
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.
A good status will look like:
status:
endpointStatus:
lastChanged: "2020-06-25T13:47:45Z"
lastChecked: "2020-06-25T13:48:46Z"
phase: Connected
A failed status:
status:
endpointStatus:
details: |-
error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
health: ERROR
This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.
Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
We can now explicitly set any property on a given pool by using the new
Property field in the CephBlockPool Spec.
Also, this fixes the case where both `CephBlockPool` and `CephCluster`
are created at the same time. When Rook creates the pool, the cluster is
still being bootstrapped and the global option
`osd_pool_default_pg_autoscale_mode` has not bee set yet. So the pool
gets created but its `pg_autoscale_mode` property is set to `warn`
instead of `on`.
Closes: https://github.com/rook/rook/issues/5608V
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit contains CR implementations for:
CephObjectRealm
CephObjectZoneGroup
CephObjectZone
Also there are changes made to the objectstore
to add rgws in the object-store to zones and
zone groups in a multisite configuration and
the removal of the --default parameter for any
realms/zonegroups/zones that are created.
Signed-off-by: Ali Maredia <amaredia@redhat.com>
This PR is for rebasing #4683 with master. With new controller runtime changes most the changes with #4683 got redundant.
Only valid change is waiting for all the object buckets to be cleaned up before removing the finalizer.
Co-authored-by: Elise Gafford <egafford@redhat.com>
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.
Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
When updating the status on resources sometimes the update fails due to
the resource being an outdated version. This frequently occurs when
the same reconcile loop updates the status multiple times, or the finalizer
is added, or some other update to the resource. Upon the next reconcile
the error would generally go away since the status or finalizer didn't
need to be updated multiple times. But now the error is not expected
since the controller will retrieve the latest version of the resource
immediately before attempting to update it.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This type was a string already and was just making us doing string()
calls all the time to it's not worth it.
Signed-off-by: Sébastien Han <seb@redhat.com>
When running on SDN, the container is not privileged and the network
stack is not exposed so running the gw on port 80 will not work (not
enough privileged).
Closes: https://github.com/rook/rook/issues/5106
Signed-off-by: Sébastien Han <seb@redhat.com>
The pools had some legacy structs that translated between the
ceph.v1 types used by the CRDs and the internal implementation
of the pools. This simplifies the pool implementation by removing
the intermediate model and leaving us only with the ceph v1
pool types and no unnecessary translation.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.
Signed-off-by: Sébastien Han <seb@redhat.com>
The pools may already be created by the mgr module. If the pool specs
are not specified in the object CR, skip pool creation for the
object store. The pools are required to exist and the reconcile
will fail until they do exist.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is critical to get this information so that the label can be set
properly and also the configuration done appropriately.
Signed-off-by: Sébastien Han <seb@redhat.com>
The bump from 0.2 to 0.4 controller-runtime versions which happened
during the use of Go Modules (bumping Go version to 1.13) introduced an
issue where the deployment annotations keep getting changed during
watch update event.
The easy fix is to stop watching Deployment objects for now.
Signed-off-by: Sébastien Han <seb@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:
* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion
Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit is to handle all those unhandled errors which raises the gosec warning.
Fixed G104: Unhandled Errors are handled now
Signed-off-by: Nizamudeen <nia@redhat.com>
When the cluster is upgraded, the child controllers are notified with the
callback to the ParentClusterChanged() method. These should be made
in a separate goroutine so that the cluster controller is not blocked
on the child controllers and also so child controllers aren't blocked
on each other for the upgrade.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We now always perform ok-to-stop change when the deployment is going to
be updated. This could be for various reasons like change in the spec or
a new Ceph image.
In any case, if Ceph resources are going to be restarted this must be
done carefully.
Also trimmed the isUpgrade variable, the upcoming commit will remove the
need of that variable.
Keeping isUpgrade from the controller code since we need it to detect
whether an orchestration upgrade should be triggered or not (based on
the ceph cluster status).
Also, we keep it on the child controllers so we don't go into an
orchestration phase if the cluster image hasn't changed. The child does
not need to be updated/go into its orchestration phase because the
cluster CR changed.
Closes: https://github.com/rook/rook/issues/4379
Signed-off-by: Sébastien Han <seb@redhat.com>