This type was a string already and was just making us doing string()
calls all the time to it's not worth it.
Signed-off-by: Sébastien Han <seb@redhat.com>
When running on SDN, the container is not privileged and the network
stack is not exposed so running the gw on port 80 will not work (not
enough privileged).
Closes: https://github.com/rook/rook/issues/5106
Signed-off-by: Sébastien Han <seb@redhat.com>
The pools had some legacy structs that translated between the
ceph.v1 types used by the CRDs and the internal implementation
of the pools. This simplifies the pool implementation by removing
the intermediate model and leaving us only with the ceph v1
pool types and no unnecessary translation.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.
Signed-off-by: Sébastien Han <seb@redhat.com>
The pools may already be created by the mgr module. If the pool specs
are not specified in the object CR, skip pool creation for the
object store. The pools are required to exist and the reconcile
will fail until they do exist.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is critical to get this information so that the label can be set
properly and also the configuration done appropriately.
Signed-off-by: Sébastien Han <seb@redhat.com>
The bump from 0.2 to 0.4 controller-runtime versions which happened
during the use of Go Modules (bumping Go version to 1.13) introduced an
issue where the deployment annotations keep getting changed during
watch update event.
The easy fix is to stop watching Deployment objects for now.
Signed-off-by: Sébastien Han <seb@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:
* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion
Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit is to handle all those unhandled errors which raises the gosec warning.
Fixed G104: Unhandled Errors are handled now
Signed-off-by: Nizamudeen <nia@redhat.com>
When the cluster is upgraded, the child controllers are notified with the
callback to the ParentClusterChanged() method. These should be made
in a separate goroutine so that the cluster controller is not blocked
on the child controllers and also so child controllers aren't blocked
on each other for the upgrade.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We now always perform ok-to-stop change when the deployment is going to
be updated. This could be for various reasons like change in the spec or
a new Ceph image.
In any case, if Ceph resources are going to be restarted this must be
done carefully.
Also trimmed the isUpgrade variable, the upcoming commit will remove the
need of that variable.
Keeping isUpgrade from the controller code since we need it to detect
whether an orchestration upgrade should be triggered or not (based on
the ceph cluster status).
Also, we keep it on the child controllers so we don't go into an
orchestration phase if the cluster image hasn't changed. The child does
not need to be updated/go into its orchestration phase because the
cluster CR changed.
Closes: https://github.com/rook/rook/issues/4379
Signed-off-by: Sébastien Han <seb@redhat.com>
The operator kit had more utility originally when the operator
was creating and managing the TPRs and CRDs directly. Since
the CRDs are now created from a manifest and no longer by the
operators, the utility of operator kit is limited to the
controller watcher. Since we are moving to the controller runtime
we simplify the code to make the transition smoother. Now
there is only a simple WatchCR method that will need to be
replaced as we maek that transition.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds status field for ceph related CRs which will be useful for knowing the ceph component status without checking the logs.
Signed-off-by: Ashish Ranjan <ashishranjan738@gmail.com>
When using the "errors" package, using `%+v` (extended format),
each Frame of the error's StackTrace will be printed in detail.
Let's only print `%v` to print the error.
If the error has a Cause it will be printed recursively.
Basically `%+v` has been replaced with `%v` for all `error` type
interfaces, whether the logger is Info, Warning or Error.
Signed-off-by: Sébastien Han <seb@redhat.com>
We recently discovered a race condition which happens not to update the
child CRDs when CR image changes.
The race happens under the following sequence:
* orchestration onAdd() is called when the operator image restarts
or its image updated
* the orchestration then goes into processing mon/mgr/osd
* during this the cluster CR is updated so onUpdate() is called,
however, onAdd() is not done yet for child CRDs (mds, rgw)
* onUpdate() runs against the cluster CR and updates the cluster images
* by the time the child CRDs reached onUpdate() from their respective
ParentClusterChanged() methods, the image is already up to date
so we were exiting since our check for updating
these resources are based on the cluster image version
To fix this, we changed the comparison, but using isUpgrade instead and
slightly changed the way isUpgrade operates. So now, isUpgrade is true
even if the Ceph version is identical which allows us to verify whether
or not the image changed.
Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1775624
Signed-off-by: Sébastien Han <seb@redhat.com>
Checking for isUpgrade is irrelevant since we need to restart the
deployment because the the spec has changed, the only thing that might
change is whether or not we perform upgrade checks.
Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1775624
Signed-off-by: Sébastien Han <seb@redhat.com>
Failure to create a keyring prior to creating a deployment
dependent on that keyring created a race conditition resulting
in an intermittent and unnecessary pod failure. This change
creates a keyring prior to any RGW deployment.
Partially fixes: #4089
Signed-off-by: egafford <egafford@redhat.com>
The resources for rgw, mds, and nfs were being cleaned up explicitly when
the CR was deleted. This is legacy from before the owner references were
set on the CRs. Now the owner references are set to the appropriate CRs so
the resources will be cleaned up upon deletion of the specific CR
rather than the cluster CR.
Signed-off-by: travisn <tnielsen@redhat.com>
We now differentiate the cases where:
* we only consume the external cluster
* we consume the external cluster as well as creating stateless
resources in Kubernetes (bootstrap mds,rgw, nfs)
This is mostly controlled via the image property spec. If not defined,
not extra CRs won't be able to be created.
Now the external cluster feature supports Ceph cluster as of Luminous 12.2.
Signed-off-by: Sébastien Han <seb@redhat.com>
Reduce the duplication of detecting the version of the external cluster
by refactoring to a single method that retrieves and validates the
version
Signed-off-by: travisn <tnielsen@redhat.com>
Add various upgrade tests so that the local cluster does not fail if
an upgrade is available on the external cluster.
Also, add better upgrades handling on update and on add. We now rely as
much as possible on a unique function that works both local and external
cluster 'detectAndValidateCephVersion()'. So the code will act the same
way wether a CR update is triggered or if the operator is restarted.
The health check now reports any topology changes on the external
cluster such as odd number of mons and low number. It will also detect
if the external cluster has been upgraded and will log that information,
helping the admin in his/her decision to trigger an image upgrade.
Finally, we prevent the creation of new CRs if the external cluster
version has a higher major number than the local one.
Closes: https://github.com/rook/rook/issues/3731
Signed-off-by: Sébastien Han <seb@redhat.com>
From now on, Rook will only perform check before upgrades when there is
an actual upgrade. So if the Ceph image changed and a new version is
desired Rook will go through all the daemons and update them one by one
and perform checks in between.
Closes: https://github.com/rook/rook/issues/3583
Signed-off-by: Sébastien Han <seb@redhat.com>
We can now deploy all the other Rook's CRDs. They will be deployed in
Kubernetes but will consume the external cluster storage.
Signed-off-by: Sébastien Han <seb@redhat.com>
When configuring an external cluster the orchestration of discover and all the ceph
daemons will be skipped. The operator will still watch for the creation of
crds for this namespace. When a filesystem, object, object user, or
ganesha crd are created, the operator will simply print an error
to the log.
Signed-off-by: travisn <tnielsen@redhat.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
When configuring an external cluster the orchestration of discover and all the ceph
daemons will be skipped. The operator will still watch for the creation of
crds for this namespace. When a filesystem, object, object user, or
ganesha crd are created, the operator will simply print an error
to the log.
Signed-off-by: travisn <tnielsen@redhat.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
rgw wasn't using UpdateDeploymentAndWait() to update its deployment
configuration. We now use it so that it's consistent with the other
daemons configuration.
Signed-off-by: Sébastien Han <seb@redhat.com>
In the 0.9 release rook converted the v1beta1 CRD resources
to v1 resources. This conversion code is no longer necessary
as we will be using v1 resources going forward.
The code paths will not be triggered anymore, thus
removing the dead code.
Signed-off-by: travisn <tnielsen@redhat.com>
The CephCluster CR contains settings that are needed by other
CRs to configure the Ceph daemons. When the CephCluster CR
is updated, the updates will now be passed on to each of the
CR controllers to ensure the daemons are updated properly
without requiring an operator restart.
When calling the controllers from another controller,
we ensure that only a single goroutine is handling
CRs at any given time to prevent contention across
multiple CRs of the same type.
Signed-off-by: travisn <tnielsen@redhat.com>
It's quite convinient to expose /var/log/ceph so that we can decide to
activate logs locally on the machine and see what's going on.
This is only a placeholder when a daemon is stuck crashlooping and we
want to allow administrator to gather log files.
We still do not log on file but this can be activated via a config
option passed to the centralized config option store.
For some daemons, which typically do not store any data (rgw, rbd-mirror
and mds) we had to propagate dataDirHostPath from the cluster spec to
each creation call so that the bindmount can happen.
Fixes: https://github.com/rook/rook/issues/2881
Signed-off-by: Sébastien Han <seb@redhat.com>
Configure the Ceph rgw daemon completely from the operator a la the
recent changes to the Ceph mon, mgr, and mds operators.
Create the rgw deployment or daemonset first, and then create the
keyring secret for the object store with its owner reference as the
corresponding deployment or daemonset. When the replication controller
is deleted, the secret is also deleted.
The RGW's mime.types file is now stored in a configmap with a different
file created for each object store. This is primarily just a means to
get the mime.types file into the rgw pod, but the added benefit is that
the administrator can modify the configmap, which could reduce
susceptibility to file type execution vulnerabilities (worst case).
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>