Commit Graph
1340 Commits
Author SHA1 Message Date
Sébastien Han 040193bb5a ceph: remove DaemonType type
This type was a string already and was just making us doing string()
calls all the time to it's not worth it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-01 09:08:18 +02:00
Sébastien Han 7788901a88 ceph: add liveness probe to mon, mds and osd daemons
Now Kubernetes will perform liveness checks on mon, mds and osd daemons.
The command will:

* call the socket (check for existence)
* execute a command and check the return code (success if 0)

This handles the case where the daemon is stuck locally and
unresponsive. It's unlikely but not impossible.
These checks bring more robustness to the implementation.

rbd-mirror and nfs have been leftover for the following reason. The
rbd-mirror socket name is different from other daemons (could be fixed
though): /run/ceph/ceph-client.rbd-mirror.a.1.94362516231272.asok also,
the command to call would need to be changed from "status" to "rbd
mirror status" so we can keep this for a later.
The nfs ganesha has no socket only a PID file which doesn't mean much.
No PID means the process does not run so Kubernetes will already handle
this and the pod will crash loop.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-31 15:20:42 +02:00
Travis Nielsen 928e901ec5 Merge pull request #5113 from leseb/rgw-port-sdn
ceph: use a different port for rgw on sdn
2020-03-30 13:32:30 -06:00
Dmitry Yusupov c2f175e266 Merge pull request #5120 from dyusupov/master
edgefs: TZ needs to be exposed
2020-03-30 12:30:35 -07:00
Dmitry Yusupov 080510bb6f edgefs: TZ needs to be exposed
/etc/timezone now exposed by operator

Signed-off-by: Dmitry Yusupov <dmitry.yusupov@nexenta.com>
2020-03-30 11:42:45 -07:00
Sébastien Han f136105951 ceph: use a different port for rgw on sdn
When running on SDN, the container is not privileged and the network
stack is not exposed so running the gw on port 80 will not work (not
enough privileged).

Closes: https://github.com/rook/rook/issues/5106
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-30 19:37:11 +02:00
Santosh Pillai 2bfd7c42bc ceph: cleanup cluster.Spec.DataDirHostPath on cluster deletion
In order to ensure proper clean up of all the rook-ceph data when the cluster is deleted, we need to clean up the dataDirHostPath (var/lib/rook)

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-03-30 19:36:01 +05:30
Elise Gafford 2ca0d8f93d ceph: add cleanupPolicy to cephcluster spec
In order to perform full automation of cephcluster deletion safely
and protect against catastrophic data loss, the end user must be
able to signify that they intend to irrecoverably delete the data
in their cluster. The cleanupPolicy field of the cluster spec is
intended to communicate this. This patch does not implement
automated deletion, but only creates the spec field and the
safety feature of halting orchestration other than deletion on a
cluster with a set cleanup policy value.

Partially-fixes: #3222
Signed-off-by: Elise Gafford <egafford@redhat.com>
2020-03-30 19:11:54 +05:30
Travis Nielsen e39637930d ceph: set pg_num for rgw metadata pools to lower default
The PG count on metadata pools should default to rgw_rados_pool_pg_num_min
instead of the more general default pg count. This means rgw pools
will default to 8 PGs instead of 32 PGs, which means a lot more pools
can be created before hitting the default PG limit.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-27 09:06:59 -06:00
Sébastien Han 9c343ae94e Merge pull request #5032 from umangachapagain/config-override
Ceph: add CSI configurations to ConfigMap
2020-03-27 15:21:24 +01:00
Umanga Chapagain 0e932c15eb Ceph: add CSI configurations to ConfigMap
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-27 15:17:30 +05:30
rohan47 3a80d7982c Ceph: Remove crashcollector pod automatically when the node is deleted.
The CrashCollector pod remains in pending state indefinitely after
running on a k8s node that was deleted. The code now deletes the
deployment after the node is deleted.

Signed-off-by: rohan47 <rohgupta@redhat.com>
2020-03-27 02:20:13 +05:30
Sébastien Han c62d89c7bc k8sutil: mitigate UpdateDeploymentAndWait condition
If a deployment stays in pending we should give early by looking at
ProgressDeadlineExceeded, this will reduce the time to wait from 20 min
to 10 min because ProgressDeadlineExceeded default is 600 seconds.

Prior to this patch we would wait 20min since we take
currentDeployment.Spec.ProgressDeadlineSeconds which is typically 600
then retry every 2 seconds, which makes it 20min total.

Closes: https://github.com/rook/rook/issues/5090
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-26 14:15:45 +01:00
Sébastien Han 55aee177f3 ceph: don't hardcode ContinueUpgradeAfterChecksEvenIfNotHealthy
Let's just read the given value from the spec.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-26 14:07:46 +01:00
Sébastien Han 80dd249ecd ceph: mds controller do not set owner ref twice
The owner ref is set right after the deployment is created one level on
top.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-26 13:00:14 +01:00
Travis Nielsen 5254de1a8c ceph: simplify pool model to v1 types
The pools had some legacy structs that translated between the
ceph.v1 types used by the CRDs and the internal implementation
of the pools. This simplifies the pool implementation by removing
the intermediate model and leaving us only with the ceph v1
pool types and no unnecessary translation.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-25 16:59:12 -06:00
Sébastien Han 8db14885b5 ceph: controller fix misleading debug log
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 19:12:17 +01:00
Sébastien Han c84d66de1b ceph: controller, reconcile faster
Let's not wait for the CephCluster to be done reconciling but instead
check for the Ceph cluster status, if it's closed to "ok" then we
proceed so HEALTH_OK and HEALTH_WARN are accepted.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 18:45:19 +01:00
Sébastien Han b5e50481ff ceph: controller: retry after 10sec even if no cluster
Even if there is no CephCluster we still want to retry every 10sec as
one will likely show up soon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Sébastien Han 7b544f04d4 ceph: reconcile more often when cluster is not ready
Let's not apply the exponential backoff when waiting for the cluster to
be ready, let's only apply it when there is no cluster.

Closes: https://github.com/rook/rook/issues/5059
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Sébastien Han a699512911 ceph: rollback to objectmatcher 1.1.0
With 1.1.1, the object matcher seems not to preserve annotation's order,
this is tracked here: https://github.com/banzaicloud/k8s-objectmatcher/issues/24

1.1.0 does not have that issue, so let's go back with 1.1.0.

Closes: https://github.com/rook/rook/issues/5047
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Travis Nielsen 006b835711 Merge pull request #5058 from travisn/object-no-pools
Allow creation of object store with pre-existing pools
2020-03-24 09:19:55 -06:00
Umanga Chapagain db3735656c Ceph: fix updates for servicemonitor
with older implementation, servicemonitor was not getting
updated due to missing resource version. This fix adds
resource version to the servicemonitor definition and
ensures that the object is properly created or
updated.

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-24 16:30:44 +05:30
Travis Nielsen 5d2db6a9f1 ceph: allow creation of object store with pre-existing pools
The pools may already be created by the mgr module. If the pool specs
are not specified in the object CR, skip pool creation for the
object store. The pools are required to exist and the reconcile
will fail until they do exist.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-23 23:54:20 -06:00
Sébastien Han f64500a954 ceph: failed reconcile if ceph version is not found
This is critical to get this information so that the label can be set
properly and also the configuration done appropriately.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-20 17:38:03 +01:00
Sébastien Han 0e2f84f998 ceph: convert NFS controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4941
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-20 17:38:03 +01:00
Sébastien Han ca0a30f38d ceph: convert Filesystem controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-19 23:34:58 +01:00
Sébastien Han 7caedefb4f Merge pull request #5039 from travisn/bulk-cleanup
Logging and exec package cleanup
2020-03-19 16:54:19 +01:00
Sébastien Han d3d4cd2f20 ceph: child controller-runtime stop watching for deployment
The bump from 0.2 to 0.4 controller-runtime versions which happened
during the use of Go Modules (bumping Go version to 1.13) introduced an
issue where the deployment annotations keep getting changed during
watch update event.
The easy fix is to stop watching Deployment objects for now.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-19 15:19:23 +01:00
Travis Nielsen 4a2ca1c1ef ceph: enable debug logging in the osd prepare job
All processes created with the exec package will now only have debug logging
for the commands and their output. The osd prepare job relies heavily
on many of these commands so we can enable debug logging for the entire
job.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen e8f9cfcb71 exec: always write commands to debug log
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen f2ecaa2bda exec: remove the unused actionName param
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Travis Nielsen 18b0e7d295 ceph: scrub ceph commands to write actions to the log
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Sébastien Han 26999b7fe4 Merge pull request #5042 from leseb/size-1-pacific
ceph: allow size option on pacific only
2020-03-18 16:37:07 +01:00
Sébastien Han 31fa13c824 Merge pull request #4530 from rajatsingh25aug/cephObjectStore
Ceph: ObjectStore replica size will not change
2020-03-18 15:18:59 +01:00
Sébastien Han 66b561e45b ceph: allow size option on pacific only
This flag only exists in Pacific, not Octopus.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-18 15:12:06 +01:00
Sébastien Han f268c897e9 ceph: Convert the Ceph ObjectStore controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Sébastien Han 711ec6095b ceph: controllers just log status error
Let's return the original error instead and just log the status change
error.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:15 -06:00
Sébastien Han b3a61bf4b9 ceph: more precise watcher object user controller
Only watch and react on resources that are matching the Kind of
CephObjectStoreUser.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:15 -06:00
RAJAT SINGH 3d6d14ea73 cephObjectstore: Replica size will not change
The pool size for cephObjectStore (object.yaml), was not getting updated in the cluster upon changing.The reason being that there was an "else" case missing to update  the replica size if the pools already existed. That will fix this issue.

Signed-off-by: RAJAT SINGH <rajasing@redhat.com>
2020-03-17 19:23:28 +05:30
moricho 7d0215d820 ci: switch to use go.mod.check target
This switches to use `go.mod.check` target instead of `go.vendor.check` target.

Signed-off-by: moricho <ikeda.morito@gmail.com>
2020-03-17 16:48:43 +09:00
moricho 437a0b1d9b go: versionup sig-storage-lib-external-provisioner/controller
version up sig-storage-lib-external-provisioner/controller to v4.1.0 and
modify related files

Signed-off-by: moricho <ikeda.morito@gmail.com>
2020-03-17 16:47:55 +09:00
moricho 8907e220ea apis: regenerate files & modify k8sutil tests
This regenerate files with code-generator/deepcopy-gen and modify the related tests.

Signed-off-by: moricho <ikeda.morito@gmail.com>
2020-03-17 16:47:53 +09:00
moricho 2beb731a7e go: move to gomodules
This switches Rook to use `go mod` instead of `dep` for
dependencies management.

Signed-off-by: moricho <ikeda.morito@gmail.com>
2020-03-17 16:47:53 +09:00
Travis Nielsen 017a560db3 Merge pull request #5021 from samkulkarni20/yb-mem-limits
YugabyteDB: Add resource limits to YugabyteDB pods
2020-03-16 11:12:07 -06:00
Sébastien Han a459a2e220 Merge pull request #5030 from leseb/no-full-err
ceph: remove dead code
2020-03-16 15:19:49 +01:00
Sébastien Han f44583b2b9 ceph: remove dead code
As part of the transition to support Ceph release **as of** Nautilus, we
left over that portion of code.
We don't need it anymore.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-16 14:01:24 +01:00
Sameer Kulkarni 909770a346 YugabyteDB: Add resource limits to YugabyteDB pods
Current YugabyteDB operator code creates Master and TServer pods without any resource requests/limits (specifically CPU and memory).
This causes the operator to run into soft/hard memory limit issue. The fix adds recommended resource requests and limits as defaults to the pods it creates.

Closes: https://github.com/yugabyte/yugabyte-db/issues/3884
Signed-off-by: Sameer Kulkarni <samkulkarni20@gmail.com>
2020-03-14 12:23:14 +05:30
Travis Nielsen 783507588e Merge pull request #5023 from leseb/allow-size-1
ceph: allow setting pool size 1 on octopus
2020-03-13 10:17:41 -06:00
Sébastien Han 99f6402563 ceph: allow setting pool size 1 on octopus
Left from https://github.com/rook/rook/pull/4895
Also more cleanup on the ceph.conf since we config is in the mon store.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-13 15:48:08 +01:00