This type was a string already and was just making us doing string()
calls all the time to it's not worth it.
Signed-off-by: Sébastien Han <seb@redhat.com>
When running on SDN, the container is not privileged and the network
stack is not exposed so running the gw on port 80 will not work (not
enough privileged).
Closes: https://github.com/rook/rook/issues/5106
Signed-off-by: Sébastien Han <seb@redhat.com>
The PG count on metadata pools should default to rgw_rados_pool_pg_num_min
instead of the more general default pg count. This means rgw pools
will default to 8 PGs instead of 32 PGs, which means a lot more pools
can be created before hitting the default PG limit.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The pools had some legacy structs that translated between the
ceph.v1 types used by the CRDs and the internal implementation
of the pools. This simplifies the pool implementation by removing
the intermediate model and leaving us only with the ceph v1
pool types and no unnecessary translation.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.
Signed-off-by: Sébastien Han <seb@redhat.com>
Let's not wait for the CephCluster to be done reconciling but instead
check for the Ceph cluster status, if it's closed to "ok" then we
proceed so HEALTH_OK and HEALTH_WARN are accepted.
Signed-off-by: Sébastien Han <seb@redhat.com>
The pools may already be created by the mgr module. If the pool specs
are not specified in the object CR, skip pool creation for the
object store. The pools are required to exist and the reconcile
will fail until they do exist.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is critical to get this information so that the label can be set
properly and also the configuration done appropriately.
Signed-off-by: Sébastien Han <seb@redhat.com>
The bump from 0.2 to 0.4 controller-runtime versions which happened
during the use of Go Modules (bumping Go version to 1.13) introduced an
issue where the deployment annotations keep getting changed during
watch update event.
The easy fix is to stop watching Deployment objects for now.
Signed-off-by: Sébastien Han <seb@redhat.com>
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
The pool size for cephObjectStore (object.yaml), was not getting updated in the cluster upon changing.The reason being that there was an "else" case missing to update the replica size if the pools already existed. That will fix this issue.
Signed-off-by: RAJAT SINGH <rajasing@redhat.com>
As part of the transition to support Ceph release **as of** Nautilus, we
left over that portion of code.
We don't need it anymore.
Signed-off-by: Sébastien Han <seb@redhat.com>
Now, the CephObjectStoreUser CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:
* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion
Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:
* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion
Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
As of Octopus, Ceph will prevent you from creating a pool with a
replica size of 1. Allowing such pool could lead to data loss, so enable
the new option: requireSafeReplicaSize: false if you are **ABSOLUTELY**
certain that is what you want.
Closes: https://github.com/rook/rook/issues/4889
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit is to handle all those unhandled errors which raises the gosec warning.
Fixed G104: Unhandled Errors are handled now
Signed-off-by: Nizamudeen <nia@redhat.com>
When the cluster is upgraded, the child controllers are notified with the
callback to the ParentClusterChanged() method. These should be made
in a separate goroutine so that the cluster controller is not blocked
on the child controllers and also so child controllers aren't blocked
on each other for the upgrade.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Versions of k8s prior to 1.14 do not support the nullable field in
CRD properties. As k8s 1.13 remains part of our support matrix and
as this field must remain nullable, we must remove CRD validation
of SecurePort. Range validation of this field has been moved into
the logic of rgw.validateStore.
Fixes: #4693
Signed-off-by: Elise Gafford <egafford@redhat.com>
Code for pod anti affinity mechanism on ceph gateways is a duplica of the one used by ceph mon.
This is a refacto to use same code on both components
Signed-off-by: n.fraison <n.fraison@criteo.com>
We now always perform ok-to-stop change when the deployment is going to
be updated. This could be for various reasons like change in the spec or
a new Ceph image.
In any case, if Ceph resources are going to be restarted this must be
done carefully.
Also trimmed the isUpgrade variable, the upcoming commit will remove the
need of that variable.
Keeping isUpgrade from the controller code since we need it to detect
whether an orchestration upgrade should be triggered or not (based on
the ceph cluster status).
Also, we keep it on the child controllers so we don't go into an
orchestration phase if the cluster image hasn't changed. The child does
not need to be updated/go into its orchestration phase because the
cluster CR changed.
Closes: https://github.com/rook/rook/issues/4379
Signed-off-by: Sébastien Han <seb@redhat.com>
A new CRD option `continueUpgradeAfterChecksEvenIfNotHealthy` is added.
When upgrading, Rook goes OSD by OSD and then waits for PGs to be clean
before proceeding to the next OSD. Currently, Rook waits for 5 hours but
there might be circumstances where PGs need more time to settle.
Thus setting `continueUpgradeAfterChecksEvenIfNotHealthy` to true will
pursue the upgrade process, even if PGs are not 100% active+clean.
Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1786029
Signed-off-by: Sébastien Han <seb@redhat.com>
The operator kit had more utility originally when the operator
was creating and managing the TPRs and CRDs directly. Since
the CRDs are now created from a manifest and no longer by the
operators, the utility of operator kit is limited to the
controller watcher. Since we are moving to the controller runtime
we simplify the code to make the transition smoother. Now
there is only a simple WatchCR method that will need to be
replaced as we maek that transition.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds status field for ceph related CRs which will be useful for knowing the ceph component status without checking the logs.
Signed-off-by: Ashish Ranjan <ashishranjan738@gmail.com>
Reduce from 5 minutes to less than 50 seconds the time needed to restart a
mgr and osds(PVCs based),rgw,mds,rbd,nfs and toolbox pod that was running
in a k8s "NotReady" Node.
A new Rook Operator setting has been added to allow the user to set the
time used in pod's toleration <node.kubernetes.io/unreachable>:
ROOK_UNREACHABLE_NODE_TOLERATION_SECONDS = <5>
New setting added in release notes
Addressed @travisn,@leseb, and @blaineEXE suggestions
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
When using the "errors" package, using `%+v` (extended format),
each Frame of the error's StackTrace will be printed in detail.
Let's only print `%v` to print the error.
If the error has a Cause it will be printed recursively.
Basically `%+v` has been replaced with `%v` for all `error` type
interfaces, whether the logger is Info, Warning or Error.
Signed-off-by: Sébastien Han <seb@redhat.com>
Both --id and --name do the same thing, so let's use one only: --id
which is shared accross all the daemons already.
Signed-off-by: Sébastien Han <seb@redhat.com>
rgw needs the fronted on its cli line to select the fronted as well the
port. This option is not supported in the mon config store.
Closes: https://github.com/rook/rook/issues/4471
Signed-off-by: Sébastien Han <seb@redhat.com>
As performed with mon, gateways from same ceph object store
must not run on the same host if hostNetwork is set to true
Signed-off-by: n.fraison <n.fraison@criteo.com>