Commit Graph
165 Commits
Author SHA1 Message Date
Redouane Kachach 0ae7867dd1 Revert "mgr: use k8s readiness probe to implement mgr HA"
This reverts commit dc76f81fea.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-03-07 13:12:50 +01:00
Redouane Kachach 0c652d28e2 Revert "ci: fix MultiClusterDeploySuite CI"
This reverts commit b464428978.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-03-07 13:12:50 +01:00
parth-gr a84daf9bf0 core: change io/ioutil package to use io and os package
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues

Signed-off-by: parth-gr <paarora@redhat.com>
2023-02-17 20:38:29 +05:30
Travis Nielsen 0a80789367 Merge pull request #11690 from rkachach/fix_issue_11685
ci: fix MultiClusterDeploySuite CI
2023-02-16 11:19:04 -07:00
Redouane Kachach b464428978 ci: fix MultiClusterDeploySuite CI
This fixes the the MultiClusterDeploySuite CI failure. The operator
is timing out waiting for the mgr deployments to be
ready, according to the WaitForDeploymentToStart() method. After
the reconcile times out after about five minutes, the next
reconcile succeeds since the wait is only done for new mgr
deployments.

Closes: https://github.com/rook/rook/issues/11685

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-02-16 18:43:38 +01:00
Travis Nielsen 3dabc6dcb6 Merge pull request #11317 from avanthakkar/introduce-ceph-exporter
core: introduce ceph-exporter
2023-02-15 11:59:05 -07:00
Travis Nielsen b008c5753f Merge pull request #11643 from rkachach/fix_issue_11640
mgr: use k8s readiness probe to implement mgr HA
2023-02-15 09:51:48 -07:00
Avan Thakkar 460900756c core: add service monitor for ceph-exporter service
Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2023-02-15 15:44:56 +05:30
Redouane Kachach dc76f81fea mgr: use k8s readiness probe to implement mgr HA
The idea behind this change is to use the readiness probe to implement
the mgr HA mechanism. In the current ceph mgr implementation only the
active instance offers the command 'mgr_status' through the admin
socket. We use this command combined with a Readiness Exec Probe to
detect which manager is active. Kubernetes will automatically mark
it as ready and redirect any service traffic to the active instance.

Closes: https://github.com/rook/rook/issues/11640
Closes: https://github.com/rook/rook/issues/11638
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-02-15 10:12:16 +01:00
subhamkrai 036715c3b2 rbdmirror: set log rotation from 7 to 4x i.e 28
increasing the rotation from default 7 to 28 as
in case of rbdmirror logs seems not enough in some cases
with maxLogSize 500 so it's better to increase the rotation
for rbdmirror specific.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-02-13 21:37:57 +05:30
Avan Thakkar 2f8ee60374 core: introduce ceph-exporter
Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2023-02-06 02:47:09 -05:00
Travis Nielsen b693a5cca5 core: refactor crash collector for more node daemons
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md

Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-12-09 16:57:05 -07:00
Travis Nielsen 3a07f3ac98 Merge pull request #10887 from travisn/duplicate-mon-endpoint
mon: Remove out of quorum mons from ceph.conf
2022-11-09 10:16:07 -07:00
Shinya Hayashi 05875a3f4f osd: support loop devices for test clusters
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.

This feature is intended to be used for testing purposes only.

Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
2022-11-09 06:45:41 +00:00
Travis Nielsen b6e8ea2b50 mon: remove out of quorum mons from ceph.conf
The mons that are out of quorum may cause ceph commands to timeout
or fail unnecessarily trying to connect to a mon that is no longer
online. Now the mon health check will update the mon endpoints configmap
when a mon is detected out of quorum. This also means that if the
operator is restarted during a mon failover, the failed mon will no
longer remain in the ceph.conf, thus allowing the quorum to be more
likely to respond to the mons that are still in quorum.

The update for mons out of quorum only applies if other mons are in
quorum. If quorum is down, the configmap will not keep track of the offline
mons since too many are offline.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-11-08 11:36:37 -07:00
Travis Nielsen df6d7af355 security: run the crash collector as ceph user
The crash collector does not have the command line arguments
to run as ceph user id 167, so we set the security context
to run as the ceph user in the main crash collector
container.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-10-27 13:12:21 -06:00
Travis Nielsen 7f0c83ad18 Merge pull request #10986 from randymtz/increase-liveness-timeout
core: increase liveness probe timeout to 5s
2022-10-17 11:49:18 -06:00
Travis Nielsen 0779618816 Merge pull request #10966 from avanthakkar/customizable-image-pull-policy
operator: make imagePullPolicy customizable for csi driver and ceph pods
2022-09-28 07:23:26 -06:00
Avan Thakkar 934aa91056 operator: make imagePullPolicy customizable for csi driver and ceph pods
Introduce a new env variable ROOK_CSI_IMAGE_PULL_POLICY in rook operator configmap which should be used to
customize the imagePullPolicy for the csi driver and imagePullPolicy property in cephVersionSpec for ceph pods.

Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2022-09-27 11:57:43 +05:30
parth-gr 26584fc6e5 core: update loadclusterInfo with multus check
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor

Signed-off-by: parth-gr <paarora@redhat.com>
2022-09-22 14:46:51 +05:30
Randy J. Martinez ac9df66b76 core: increase liveness probe timeout to 5s
stability issues have been observed with 1s.
socket latency is expected whenever CPUs are
under minor pressure. Increasing value to
5s should cover most small-medium scale envs.

Resolves BZ: 2126566

Signed-off-by: Randy J. Martinez <randy@cephtips.com>
2022-09-13 17:32:43 -05:00
motorailgun ff0951738f operator: improve ProbeHandler error message
This commit implements more diagnostic error for unsccessful
Liveness- and Readiness-Probe of Pods.

Added codes are expected to catch failures of
`ceph status` and `ceph mon_status`, and report it.

Closes: https://github.com/rook/rook/issues/9846
Closes: https://github.com/rook/rook/issues/9852

Signed-off-by: motorailgun <motoi_public@mail.aria-on-the-planet.es>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-09-07 07:35:33 +00:00
subhamkrai 9d4d5bc702 core: fix logrotate bash check and periodicity logic
we need use `!=` for string comparision in bash instead of
`-ne`. Also, need to correct periodicity if condition to
make it work as expected.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-08-22 17:14:51 +05:30
subhamkrai 62f73dcd98 core: add support to rotate log based on logfile size
this commits add new field `MaxLogSize` inside `LogCollectorSpec` of
cephCluster cr which will take max size of log after which we want to
rotate the log.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-07-26 13:06:39 +00:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Travis Nielsen dad97f3425 core: remove support for ceph octopus
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-07-07 15:03:26 -06:00
subhamkrai 62f0fb2b1d core: increase liveness probe timeout to 2s
we have noticed multiple failures because of the probe
failing, most of the time it's due to fewer resources.
But increasing timeout fixes that, so increasing the
probe timeout to 2s from default 1s so that it will
give more time to probe before failing.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-06-17 19:06:40 +05:30
Travis Nielsen 28e721d877 osd: allow the osd to take a long time to start
The startup probe for the OSD has been too aggressive to kill the OSD
in case the OSD is taking some time to start. The OSD may be self-optimizing,
scrubbing, or some other internal operation before it is ready to start.
Rather than disable the startup probe completely, the default is now
to retry for two hours in case the OSD is performing those operations.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-05-11 14:29:25 -06:00
Sébastien Han 583791c45c core: move clusterInfo code to the controller package
The CSI package needs to load clusterInfo, today this code is in the mon
package which makes the call of LoadClusterInfo impossible without
having a circular import.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 11:05:02 +02:00
Alexander Trost 8686296e17 core: remove double imported packages
This removes double package imports. Example:
```
"github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
cephv1 "github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
```
Only one is now being used as shown in go-staticcheck ST1019

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-25 13:51:45 +02:00
subhamkrai bf7daccf60 core: make code changes to support latest cntrl runtime
making necessary code changes to support controller
runtime version.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-20 22:30:15 +05:30
Madhu Rajanna b0fc7c9b92 namespace: add new CRD
This introduces a new CRD to add the ability
to create rados namespace for a given
ceph block pool. Typically the name of the pool
is the name of the blockpool created by rook.

Closes: #7035

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-04-05 10:10:04 +05:30
subhamkrai 24802c559e core: fix golangci linter
fix golangci linter

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-04 20:59:31 +05:30
Sébastien Han 05506e7a68 core: reload go routine after CR is edited
Previously, the struct maintaining the list of cluster was still
initialized with a cluster item. Then the monitoring check will see that
the cluster is part of the struct already and thus won't run the
monitoring go routine again.
Now each time we cancel the context, we also remove the cluster item
from the map so that when the controller runs again, the monitoring
struct is re-populated and the go routine runs and statuses are updated.

Closes: https://github.com/rook/rook/issues/9911
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-03-31 08:49:42 +02:00
Travis Nielsen 4edcff04f7 monitoring: create prometheus rules with helm chart
The prometheus rules had been previously created if the cephcluster CR
setting monitoring.enabled was set to true. The rules were not customizable
and therefore not flexible enough. Now the rules are installed by the helm
chart. To customize the rules, a post-processor can be applied to the helm
chart.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-03-21 14:47:52 -06:00
parth-gr 2dfd64a97c core: add observedGeneration to CR status
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller

Closes: https://github.com/rook/rook/issues/9673

Signed-off-by: parth-gr <paarora@redhat.com>
2022-03-16 19:58:35 +05:30
Sébastien Han 5404ec13a2 core: dereference pointer before trying to compare with deepequal
Prior to this, we were comparing a pointer (the memory address) with a
struct. This was obviously always failing and returned false. We must
dereference the pointer to access the data contained at that memory
location.

Closes: https://github.com/rook/rook/issues/9544
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-27 17:41:42 +01:00
Sébastien Han 1208bc1410 core: reconcile operator configuration with env var
Previously, we were ignoring configuration changes coming from the
operator's pod env variables. We were assuming most users were using the
operator configmap "rook-ceph-operator-config" but most Helm users
don't. Now we will reconcile if a cephcluster is found during a CREATE
event (can be an operator restart or a cephcluster creation) AND no
"rook-ceph-operator-config" is found which means the operator's pod env
var are used.
Also, unit tests have been added (long due) for the predicate!

Closes: #9602, #9487, #9579
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-20 16:13:11 +01:00
Blaine Gardner 5b9ac16670 Merge pull request #9468 from BlaineEXE/startup-probes
core: rgw: allow specifying daemon startup probes
2022-01-04 11:03:12 -07:00
Blaine Gardner c07d89d9ea core: rgw: allow specifying daemon startup probes
Allow specifying daemon startup probes where we also allow configuring
liveness probes. Startup probes allow Rook to tolerate when Ceph daemons
occasionally take a long time to start up while not also making
Kubernetes liveness probes slower to detect runtime failures of daemons.

Startup probes are beta in Kubernetes 1.18, so we should not enable
probes by default for earlier Kubernetes versions.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-21 15:12:41 -07:00
Blaine Gardner bb0d0d39f6 Merge pull request #9384 from leseb/fix-7036
subvolumegroup: add new crd
2021-12-21 10:24:23 -07:00
Sébastien Han 6e9eb33782 subvolumegroup: add new crd
This introduces a new CRD to add the ability to create subvolumegroup
for a given ceph filesystem volume. Typically the name of the volume is
the name of the filesystem created by rook.

Closes: https://github.com/rook/rook/issues/7036
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-21 16:28:38 +01:00
Sébastien Han fffb862956 Merge pull request #9457 from leseb/fix-9452
core: disallow multiple clusters in the same namespace
2021-12-20 17:18:10 +01:00
Sébastien Han da76a772f8 core: disallow multiple clusters in the same namespace
Rook does not support running multiple clusters in the same namespace,
so the operator should not reconcile if a new cluster is added.
The scenario where a cluster is added while the operator is down is also
handled. CR updates are also handled.
When the operator detects more than one cluster it will refuse to
reconcile the CephCluster, and child CRDs will block too until the
operator is ready.
The user must remove one of the clusters before can continue to perform
any reconcile.

Closes: https://github.com/rook/rook/issues/9452
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-17 10:56:15 +01:00
Travis Nielsen 7ee9cc9d56 pool: allow configuration of built-in pools with non-k8s names
The built-in pools device_health_metrics and .nfs created by ceph
need to be configured for replicas, failure domain, etc.
To support this, we allow the pool to be created as a CR.
Since K8s does not support underscores in the resource names
the operator must translate this special pool name into
the name expected by ceph.

This also sets the basis for allowing filesystem data
pools to specify the desired pool name instead of requiring
a generated name.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-15 15:32:02 -07:00
Sébastien Han 5c1e459a6a mgr: run active-watch as root and privileged
For now, we must run the container with UID 0 and privileged for
multiple reasons:

* the rook binary writes ceph config to /var/lib/rook which is owned by
  root
* it's difficult to use /etc/ceph since it will conflict with the
  rook-ceph-override configmap AND is also owned by root since it's a
  mounted configmap.
* using /etc/ceph might be possible but has other issues with rook's
  exec package since the ceph config is built from /var/lib/rook

Closes: https://github.com/rook/rook/issues/9385
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-13 14:02:33 +01:00
Sébastien Han 870114fc91 Merge pull request #9393 from y1r/add-context-k8sutil-endpoint
core: add context parameter to k8sutil endpoint
2021-12-13 10:04:19 +01:00
Yuichiro Ueno b6ce262e7f core: add context to k8sutil replicaset and secret
This commit adds context parameter to k8sutil replicaset and secret
functions. By this, we can handle cancellation during API call of
replicaset and secret resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-12-11 11:42:31 +09:00
Yuichiro Ueno d1d252c5c2 core: add context parameter to k8sutil endpoint
This commit adds context parameter to k8sutil endpoint functions. By
this, we can handle cancellation during API call of endpoint resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-12-11 11:26:21 +09:00
Blaine Gardner a99f8985b0 Merge pull request #9150 from yuvalif/ceph-notifications-refactor
ceph: bucket notifications refactoring work
2021-12-08 11:30:15 -07:00