Commit Graph
133 Commits
Author SHA1 Message Date
Sébastien Han 040193bb5a ceph: remove DaemonType type
This type was a string already and was just making us doing string()
calls all the time to it's not worth it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-01 09:08:18 +02:00
Sébastien Han f136105951 ceph: use a different port for rgw on sdn
When running on SDN, the container is not privileged and the network
stack is not exposed so running the gw on port 80 will not work (not
enough privileged).

Closes: https://github.com/rook/rook/issues/5106
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-30 19:37:11 +02:00
Travis Nielsen e39637930d ceph: set pg_num for rgw metadata pools to lower default
The PG count on metadata pools should default to rgw_rados_pool_pg_num_min
instead of the more general default pg count. This means rgw pools
will default to 8 PGs instead of 32 PGs, which means a lot more pools
can be created before hitting the default PG limit.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-27 09:06:59 -06:00
Sébastien Han 55aee177f3 ceph: don't hardcode ContinueUpgradeAfterChecksEvenIfNotHealthy
Let's just read the given value from the spec.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-26 14:07:46 +01:00
Travis Nielsen 5254de1a8c ceph: simplify pool model to v1 types
The pools had some legacy structs that translated between the
ceph.v1 types used by the CRDs and the internal implementation
of the pools. This simplifies the pool implementation by removing
the intermediate model and leaving us only with the ceph v1
pool types and no unnecessary translation.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-25 16:59:12 -06:00
Sébastien Han 8db14885b5 ceph: controller fix misleading debug log
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 19:12:17 +01:00
Sébastien Han c84d66de1b ceph: controller, reconcile faster
Let's not wait for the CephCluster to be done reconciling but instead
check for the Ceph cluster status, if it's closed to "ok" then we
proceed so HEALTH_OK and HEALTH_WARN are accepted.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 18:45:19 +01:00
Sébastien Han a699512911 ceph: rollback to objectmatcher 1.1.0
With 1.1.1, the object matcher seems not to preserve annotation's order,
this is tracked here: https://github.com/banzaicloud/k8s-objectmatcher/issues/24

1.1.0 does not have that issue, so let's go back with 1.1.0.

Closes: https://github.com/rook/rook/issues/5047
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Travis Nielsen 5d2db6a9f1 ceph: allow creation of object store with pre-existing pools
The pools may already be created by the mgr module. If the pool specs
are not specified in the object CR, skip pool creation for the
object store. The pools are required to exist and the reconcile
will fail until they do exist.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-23 23:54:20 -06:00
Sébastien Han f64500a954 ceph: failed reconcile if ceph version is not found
This is critical to get this information so that the label can be set
properly and also the configuration done appropriately.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-20 17:38:03 +01:00
Sébastien Han 7caedefb4f Merge pull request #5039 from travisn/bulk-cleanup
Logging and exec package cleanup
2020-03-19 16:54:19 +01:00
Sébastien Han d3d4cd2f20 ceph: child controller-runtime stop watching for deployment
The bump from 0.2 to 0.4 controller-runtime versions which happened
during the use of Go Modules (bumping Go version to 1.13) introduced an
issue where the deployment annotations keep getting changed during
watch update event.
The easy fix is to stop watching Deployment objects for now.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-19 15:19:23 +01:00
Travis Nielsen e8f9cfcb71 exec: always write commands to debug log
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen f2ecaa2bda exec: remove the unused actionName param
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Travis Nielsen 18b0e7d295 ceph: scrub ceph commands to write actions to the log
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Sébastien Han 31fa13c824 Merge pull request #4530 from rajatsingh25aug/cephObjectStore
Ceph: ObjectStore replica size will not change
2020-03-18 15:18:59 +01:00
Sébastien Han f268c897e9 ceph: Convert the Ceph ObjectStore controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Sébastien Han 711ec6095b ceph: controllers just log status error
Let's return the original error instead and just log the status change
error.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:15 -06:00
Sébastien Han b3a61bf4b9 ceph: more precise watcher object user controller
Only watch and react on resources that are matching the Kind of
CephObjectStoreUser.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:15 -06:00
RAJAT SINGH 3d6d14ea73 cephObjectstore: Replica size will not change
The pool size for cephObjectStore (object.yaml), was not getting updated in the cluster upon changing.The reason being that there was an "else" case missing to update  the replica size if the pools already existed. That will fix this issue.

Signed-off-by: RAJAT SINGH <rajasing@redhat.com>
2020-03-17 19:23:28 +05:30
moricho 2beb731a7e go: move to gomodules
This switches Rook to use `go mod` instead of `dep` for
dependencies management.

Signed-off-by: moricho <ikeda.morito@gmail.com>
2020-03-17 16:47:53 +09:00
Sébastien Han f44583b2b9 ceph: remove dead code
As part of the transition to support Ceph release **as of** Nautilus, we
left over that portion of code.
We don't need it anymore.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-16 14:01:24 +01:00
Sébastien Han 9f2867e12a ceph: separate controller for CephObjectStoreUser CRD
Now, the CephObjectStoreUser CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-06 11:53:40 +01:00
Sébastien Han a3068dee0b ceph: separate controller for CephBlockPool CRD
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-02 17:31:03 +01:00
Sébastien Han dab9233c8b ceph: add CRD setting for pool size 1
As of Octopus, Ceph will prevent you from creating a pool with a
replica size of 1. Allowing such pool could lead to data loss, so enable
the new option: requireSafeReplicaSize: false if you are **ABSOLUTELY**
certain that is what you want.

Closes: https://github.com/rook/rook/issues/4889
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-02-25 16:28:58 +01:00
Nizamudeen 53883f68cf ceph: Handling Unhandled errors
This commit is to handle all those unhandled errors which raises the gosec warning.

Fixed G104: Unhandled Errors are handled now

Signed-off-by: Nizamudeen <nia@redhat.com>
2020-02-21 22:48:24 +05:30
Travis Nielsen 41fe8c3a2b ceph: call the ParentClusterChanged in a goroutine
When the cluster is upgraded, the child controllers are notified with the
callback to the ParentClusterChanged() method. These should be made
in a separate goroutine so that the cluster controller is not blocked
on the child controllers and also so child controllers aren't blocked
on each other for the upgrade.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-02-17 16:24:28 -07:00
Sébastien Han 45118185cf ceph: remove mimic support, default to nautilus
As of 1.3, Rook will only support Ceph Nautilus 14.2.5.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-02-13 09:17:34 +01:00
Travis Nielsen d40cb217ae Merge pull request #4694 from egafford/secure-port-for-ceph-object-manifests
ceph: add securePort values for objectstore integration tests
2020-01-27 14:49:47 -07:00
Elise Gafford d7906266e1 ceph: remove validation on SecurePort for k8s <= 1.13
Versions of k8s prior to 1.14 do not support the nullable field in
CRD properties. As k8s 1.13 remains part of our support matrix and
as this field must remain nullable, we must remove CRD validation
of SecurePort. Range validation of this field has been moved into
the logic of rgw.validateStore.

Fixes: #4693
Signed-off-by: Elise Gafford <egafford@redhat.com>
2020-01-27 13:49:35 -05:00
Sébastien Han d7a9409a9e Merge pull request #4606 from ashangit/set_pod_placement
ceph: create a generic setPodPlacement used by mon and gateways
2020-01-27 14:58:59 +01:00
n.fraison 383299c5f5 k8sutil: refacto pod anti affinity placement used by ceph mon and gateways
Code for pod anti affinity mechanism on ceph gateways is a duplica of the one used by ceph mon.
This is a refacto to use same code on both components

Signed-off-by: n.fraison <n.fraison@criteo.com>
2020-01-27 09:58:50 +01:00
Sébastien Han caa8382e25 ceph: do not update deployment if no changes
We know use k8s-objectmatcher to verify whether a deployment is going to
change or not.
This avoids calling Update() from k8s utils which takes time to run.

Closes: https://github.com/rook/rook/issues/4519 and https://github.com/rook/rook/issues/4642
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-24 17:02:40 +01:00
Sébastien Han 92c1696bfc ceph: cleanup/trim isUpgrade variable a bit
We now always perform ok-to-stop change when the deployment is going to
be updated. This could be for various reasons like change in the spec or
a new Ceph image.
In any case, if Ceph resources are going to be restarted this must be
done carefully.

Also trimmed the isUpgrade variable, the upcoming commit will remove the
need of that variable.

Keeping isUpgrade from the controller code since we need it to detect
whether an orchestration upgrade should be triggered or not (based on
the ceph cluster status).

Also, we keep it on the child controllers so we don't go into an
orchestration phase if the cluster image hasn't changed. The child does
not need to be updated/go into its orchestration phase because the
cluster CR changed.

Closes: https://github.com/rook/rook/issues/4379
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-24 17:02:37 +01:00
Sébastien Han 05aa639317 ceph: add option to continue on unclean PGs
A new CRD option `continueUpgradeAfterChecksEvenIfNotHealthy` is added.
When upgrading, Rook goes OSD by OSD and then waits for PGs to be clean
before proceeding to the next OSD. Currently, Rook waits for 5 hours but
there might be circumstances where PGs need more time to settle.
Thus setting `continueUpgradeAfterChecksEvenIfNotHealthy` to true will
pursue the upgrade process, even if PGs are not 100% active+clean.

Closes: https://bugzilla.redhat.com/show_bug.cgi?id=1786029
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-01-07 12:18:54 -07:00
Travis Nielsen d0c7d28a63 build: remove operator kit dependency
The operator kit had more utility originally when the operator
was creating and managing the TPRs and CRDs directly. Since
the CRDs are now created from a manifest and no longer by the
operators, the utility of operator kit is limited to the
controller watcher. Since we are moving to the controller runtime
we simplify the code to make the transition smoother. Now
there is only a simple WatchCR method that will need to be
replaced as we maek that transition.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-01-07 08:30:14 -07:00
Sébastien Han 805e606cce Merge pull request #4508 from bsperduto/hotfix/objectBucketClaimPort
ceph: fix object bucket provisioner when rgw not on port 80
2019-12-20 10:37:23 +01:00
Travis Nielsen f48e0948ee Merge pull request #4355 from ashishranjan738/status
enhance(ceph): Adds status field for ceph related CRs
2019-12-17 11:57:20 -07:00
Ashish Ranjan 8274aa2426 enhance(ceph): Adds status field for ceph related CRs
This commit adds status field for ceph related CRs which will be useful for knowing the ceph component status without checking the logs.

Signed-off-by: Ashish Ranjan <ashishranjan738@gmail.com>
2019-12-17 12:01:17 +05:30
Juan Miguel Olmo Martínez e8ccef854d Ceph: Toleration seconds for <NotReady> nodes is now configurable
Reduce from 5 minutes to less than 50 seconds the time needed to restart a
mgr and osds(PVCs based),rgw,mds,rbd,nfs and toolbox pod that was running
in a k8s "NotReady" Node.

A new Rook Operator setting has been added to allow the user to set the
time used in pod's toleration <node.kubernetes.io/unreachable>:
ROOK_UNREACHABLE_NODE_TOLERATION_SECONDS = <5>

New setting added in release notes

Addressed @travisn,@leseb, and @blaineEXE suggestions

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2019-12-16 21:38:46 -07:00
Brian Sperduto 78f72d48d6 ceph: fix object bucket provisioner when rgw not on port 80
Solves issue where provisioner would not handle other ports

Signed-off-by: Brian Sperduto <brian.sperduto@gmail.com>
2019-12-16 19:29:21 -06:00
Sébastien Han ad95c7296f ceph: do not print extended format for loggers.
When using the "errors" package, using `%+v` (extended format),
each Frame of the error's StackTrace will be printed in detail.
Let's only print `%v` to print the error.
If the error has a Cause it will be printed recursively.

Basically `%+v` has been replaced with `%v` for all `error` type
interfaces, whether the logger is Info, Warning or Error.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-16 18:35:54 +01:00
Sébastien Han 991c8d6e4d Merge pull request #4484 from mkogan1/rgw-healthcheck
ceph: rgw change the liveness probe url
2019-12-12 18:06:43 +01:00
Mark Kogan 3c807688fe ceph: rgw change the liveness probe url
to the RGWs swift objectstorage healthcheck api url [1]
which resolves an OCS false-positive CLBO of the RGW pod.

[1] https://docs.openstack.org/swift/latest/middleware.html#healthcheck

Signed-off-by: Mark Kogan <mkogan@redhat.com>
2019-12-12 12:47:11 +02:00
Blaine Gardner adfc1ef9a9 Merge pull request #4457 from ashangit/gateways_podantiaffinity
ceph: enforce pod anti affinity on gateways if hostNetwork is true
2019-12-11 12:02:42 -07:00
Sébastien Han 9cb752a874 ceph: rgw fix --id and remove --name
Both --id and --name do the same thing, so let's use one only: --id
which is shared accross all the daemons already.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-11 15:24:42 +01:00
Sébastien Han dde9e189b4 ceph: rgw fix fronted flag
rgw needs the fronted on its cli line to select the fronted as well the
port. This option is not supported in the mon config store.

Closes: https://github.com/rook/rook/issues/4471
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-11 15:24:26 +01:00
n.fraison b9c2d616ab ceph: enforce pod anti affinity on gateways if hostNetwork is true
As performed with mon, gateways from same ceph object store
must not run on the same host if hostNetwork is set to true

Signed-off-by: n.fraison <n.fraison@criteo.com>
2019-12-11 14:53:07 +01:00
Sébastien Han 5ce2ed220e ceph: use "github.com/pkg/errors"
We now use the error package.
Kubernetes errors have been renamed kerrors since they are lower than
'errors'.

Closes: https://github.com/rook/rook/issues/4054
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-12-09 16:58:32 +01:00
Mateusz Los 0daec4a7d6 ceph: change rgw client prefix
add rgw prefix to objectstore user

Signed-off-by: Mateusz Los <los.mateusz@gmail.com>
2019-12-09 12:49:24 +01:00