Commit Graph
150 Commits
Author SHA1 Message Date
Travis Nielsen df6d7af355 security: run the crash collector as ceph user
The crash collector does not have the command line arguments
to run as ceph user id 167, so we set the security context
to run as the ceph user in the main crash collector
container.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-10-27 13:12:21 -06:00
Travis Nielsen 7f0c83ad18 Merge pull request #10986 from randymtz/increase-liveness-timeout
core: increase liveness probe timeout to 5s
2022-10-17 11:49:18 -06:00
Travis Nielsen 0779618816 Merge pull request #10966 from avanthakkar/customizable-image-pull-policy
operator: make imagePullPolicy customizable for csi driver and ceph pods
2022-09-28 07:23:26 -06:00
Avan Thakkar 934aa91056 operator: make imagePullPolicy customizable for csi driver and ceph pods
Introduce a new env variable ROOK_CSI_IMAGE_PULL_POLICY in rook operator configmap which should be used to
customize the imagePullPolicy for the csi driver and imagePullPolicy property in cephVersionSpec for ceph pods.

Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2022-09-27 11:57:43 +05:30
parth-gr 26584fc6e5 core: update loadclusterInfo with multus check
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor

Signed-off-by: parth-gr <paarora@redhat.com>
2022-09-22 14:46:51 +05:30
Randy J. Martinez ac9df66b76 core: increase liveness probe timeout to 5s
stability issues have been observed with 1s.
socket latency is expected whenever CPUs are
under minor pressure. Increasing value to
5s should cover most small-medium scale envs.

Resolves BZ: 2126566

Signed-off-by: Randy J. Martinez <randy@cephtips.com>
2022-09-13 17:32:43 -05:00
motorailgun ff0951738f operator: improve ProbeHandler error message
This commit implements more diagnostic error for unsccessful
Liveness- and Readiness-Probe of Pods.

Added codes are expected to catch failures of
`ceph status` and `ceph mon_status`, and report it.

Closes: https://github.com/rook/rook/issues/9846
Closes: https://github.com/rook/rook/issues/9852

Signed-off-by: motorailgun <motoi_public@mail.aria-on-the-planet.es>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-09-07 07:35:33 +00:00
subhamkrai 9d4d5bc702 core: fix logrotate bash check and periodicity logic
we need use `!=` for string comparision in bash instead of
`-ne`. Also, need to correct periodicity if condition to
make it work as expected.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-08-22 17:14:51 +05:30
subhamkrai 62f73dcd98 core: add support to rotate log based on logfile size
this commits add new field `MaxLogSize` inside `LogCollectorSpec` of
cephCluster cr which will take max size of log after which we want to
rotate the log.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-07-26 13:06:39 +00:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Travis Nielsen dad97f3425 core: remove support for ceph octopus
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-07-07 15:03:26 -06:00
subhamkrai 62f0fb2b1d core: increase liveness probe timeout to 2s
we have noticed multiple failures because of the probe
failing, most of the time it's due to fewer resources.
But increasing timeout fixes that, so increasing the
probe timeout to 2s from default 1s so that it will
give more time to probe before failing.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-06-17 19:06:40 +05:30
Travis Nielsen 28e721d877 osd: allow the osd to take a long time to start
The startup probe for the OSD has been too aggressive to kill the OSD
in case the OSD is taking some time to start. The OSD may be self-optimizing,
scrubbing, or some other internal operation before it is ready to start.
Rather than disable the startup probe completely, the default is now
to retry for two hours in case the OSD is performing those operations.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-05-11 14:29:25 -06:00
Sébastien Han 583791c45c core: move clusterInfo code to the controller package
The CSI package needs to load clusterInfo, today this code is in the mon
package which makes the call of LoadClusterInfo impossible without
having a circular import.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 11:05:02 +02:00
Alexander Trost 8686296e17 core: remove double imported packages
This removes double package imports. Example:
```
"github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
cephv1 "github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
```
Only one is now being used as shown in go-staticcheck ST1019

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-25 13:51:45 +02:00
subhamkrai bf7daccf60 core: make code changes to support latest cntrl runtime
making necessary code changes to support controller
runtime version.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-20 22:30:15 +05:30
Madhu Rajanna b0fc7c9b92 namespace: add new CRD
This introduces a new CRD to add the ability
to create rados namespace for a given
ceph block pool. Typically the name of the pool
is the name of the blockpool created by rook.

Closes: #7035

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-04-05 10:10:04 +05:30
subhamkrai 24802c559e core: fix golangci linter
fix golangci linter

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-04 20:59:31 +05:30
Sébastien Han 05506e7a68 core: reload go routine after CR is edited
Previously, the struct maintaining the list of cluster was still
initialized with a cluster item. Then the monitoring check will see that
the cluster is part of the struct already and thus won't run the
monitoring go routine again.
Now each time we cancel the context, we also remove the cluster item
from the map so that when the controller runs again, the monitoring
struct is re-populated and the go routine runs and statuses are updated.

Closes: https://github.com/rook/rook/issues/9911
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-03-31 08:49:42 +02:00
Travis Nielsen 4edcff04f7 monitoring: create prometheus rules with helm chart
The prometheus rules had been previously created if the cephcluster CR
setting monitoring.enabled was set to true. The rules were not customizable
and therefore not flexible enough. Now the rules are installed by the helm
chart. To customize the rules, a post-processor can be applied to the helm
chart.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-03-21 14:47:52 -06:00
parth-gr 2dfd64a97c core: add observedGeneration to CR status
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller

Closes: https://github.com/rook/rook/issues/9673

Signed-off-by: parth-gr <paarora@redhat.com>
2022-03-16 19:58:35 +05:30
Sébastien Han 5404ec13a2 core: dereference pointer before trying to compare with deepequal
Prior to this, we were comparing a pointer (the memory address) with a
struct. This was obviously always failing and returned false. We must
dereference the pointer to access the data contained at that memory
location.

Closes: https://github.com/rook/rook/issues/9544
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-27 17:41:42 +01:00
Sébastien Han 1208bc1410 core: reconcile operator configuration with env var
Previously, we were ignoring configuration changes coming from the
operator's pod env variables. We were assuming most users were using the
operator configmap "rook-ceph-operator-config" but most Helm users
don't. Now we will reconcile if a cephcluster is found during a CREATE
event (can be an operator restart or a cephcluster creation) AND no
"rook-ceph-operator-config" is found which means the operator's pod env
var are used.
Also, unit tests have been added (long due) for the predicate!

Closes: #9602, #9487, #9579
Signed-off-by: Sébastien Han <seb@redhat.com>
2022-01-20 16:13:11 +01:00
Blaine Gardner 5b9ac16670 Merge pull request #9468 from BlaineEXE/startup-probes
core: rgw: allow specifying daemon startup probes
2022-01-04 11:03:12 -07:00
Blaine Gardner c07d89d9ea core: rgw: allow specifying daemon startup probes
Allow specifying daemon startup probes where we also allow configuring
liveness probes. Startup probes allow Rook to tolerate when Ceph daemons
occasionally take a long time to start up while not also making
Kubernetes liveness probes slower to detect runtime failures of daemons.

Startup probes are beta in Kubernetes 1.18, so we should not enable
probes by default for earlier Kubernetes versions.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-21 15:12:41 -07:00
Blaine Gardner bb0d0d39f6 Merge pull request #9384 from leseb/fix-7036
subvolumegroup: add new crd
2021-12-21 10:24:23 -07:00
Sébastien Han 6e9eb33782 subvolumegroup: add new crd
This introduces a new CRD to add the ability to create subvolumegroup
for a given ceph filesystem volume. Typically the name of the volume is
the name of the filesystem created by rook.

Closes: https://github.com/rook/rook/issues/7036
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-21 16:28:38 +01:00
Sébastien Han fffb862956 Merge pull request #9457 from leseb/fix-9452
core: disallow multiple clusters in the same namespace
2021-12-20 17:18:10 +01:00
Sébastien Han da76a772f8 core: disallow multiple clusters in the same namespace
Rook does not support running multiple clusters in the same namespace,
so the operator should not reconcile if a new cluster is added.
The scenario where a cluster is added while the operator is down is also
handled. CR updates are also handled.
When the operator detects more than one cluster it will refuse to
reconcile the CephCluster, and child CRDs will block too until the
operator is ready.
The user must remove one of the clusters before can continue to perform
any reconcile.

Closes: https://github.com/rook/rook/issues/9452
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-17 10:56:15 +01:00
Travis Nielsen 7ee9cc9d56 pool: allow configuration of built-in pools with non-k8s names
The built-in pools device_health_metrics and .nfs created by ceph
need to be configured for replicas, failure domain, etc.
To support this, we allow the pool to be created as a CR.
Since K8s does not support underscores in the resource names
the operator must translate this special pool name into
the name expected by ceph.

This also sets the basis for allowing filesystem data
pools to specify the desired pool name instead of requiring
a generated name.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-15 15:32:02 -07:00
Sébastien Han 5c1e459a6a mgr: run active-watch as root and privileged
For now, we must run the container with UID 0 and privileged for
multiple reasons:

* the rook binary writes ceph config to /var/lib/rook which is owned by
  root
* it's difficult to use /etc/ceph since it will conflict with the
  rook-ceph-override configmap AND is also owned by root since it's a
  mounted configmap.
* using /etc/ceph might be possible but has other issues with rook's
  exec package since the ceph config is built from /var/lib/rook

Closes: https://github.com/rook/rook/issues/9385
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-12-13 14:02:33 +01:00
Sébastien Han 870114fc91 Merge pull request #9393 from y1r/add-context-k8sutil-endpoint
core: add context parameter to k8sutil endpoint
2021-12-13 10:04:19 +01:00
Yuichiro Ueno b6ce262e7f core: add context to k8sutil replicaset and secret
This commit adds context parameter to k8sutil replicaset and secret
functions. By this, we can handle cancellation during API call of
replicaset and secret resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-12-11 11:42:31 +09:00
Yuichiro Ueno d1d252c5c2 core: add context parameter to k8sutil endpoint
This commit adds context parameter to k8sutil endpoint functions. By
this, we can handle cancellation during API call of endpoint resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-12-11 11:26:21 +09:00
Blaine Gardner a99f8985b0 Merge pull request #9150 from yuvalif/ceph-notifications-refactor
ceph: bucket notifications refactoring work
2021-12-08 11:30:15 -07:00
Yuval Lifshitz f2b065a3ae rgw: refactor bucket notification code
this should make sure code reuse between the OBC label controller and
the CephBucketNotification controller.
done as part of the cleanup work from:
https://github.com/rook/rook/pull/8426

this also include adding unit tests for:
- topic controller
- notification controller
- obc label controller

Signed-off-by: Yuval Lifshitz <ylifshit@redhat.com>
2021-12-07 18:24:46 +02:00
parth-gr 0a86d26b2e core: create rook resources with k8s recommended labels
Adding Recommended Labels on the resources created by rook
    and using Recommended Labels in the helm chart,
    for better visuals and management of k8s object

Closes: https://github.com/rook/rook/issues/8400
Signed-off-by: parth-gr <paarora@redhat.com>
2021-12-07 18:16:44 +05:30
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Omar Pakker 8f9055809f osd: add privileged support (back) to blkdevmapper securityContext (work-around)
The blockdevmapper securityContext was changed to request a minimal set of
required capabilities for its operation and drop running as privileged.
While the base change works and is valid in terms of the container's copy operation,
it turns out that OpenShift may require some additional configuration not
currently covered by the limited securityContext and the capabilities granted.

To not break those OpenShift deployments, make the blkdevmapper securityContext
listen to the ROOK_HOSTPATH_REQUIRES_PRIVILEGED flag again to set privileged mode.
This flag is true on OpenShift deployments and running as privileged
works around the (missing) configuration problem for now.
To properly drop privileged completely some additional investigation needs
to be done on OpenShift deployments without relying on privileged execution.

Signed-off-by: Omar Pakker <Omar007@users.noreply.github.com>
2021-11-17 12:25:13 +01:00
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
Yuval Lifshitz 71ed45b69b rgw: implement bucket notifications for object storage
following the design from here:
https://github.com/rook/rook/blob/master/design/ceph/object/ceph-bucket-notification-crd.md

Closes: https://github.com/rook/rook/issues/5313
Signed-off-by: Yuval Lifshitz <ylifshit@redhat.com>
2021-11-04 11:20:40 +02:00
Travis Nielsen fd10d98dc6 core: treat cluster as not existing if the cleanup policy is set
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-27 10:25:06 -06:00
Yuichiro Ueno 3fd86f83ae core: add context parameter to opcontroller
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-10-25 20:45:06 +09:00
parth-gr 7c99858a77 ceph: add finalizers to rook-ceph-mon secrets and configmap
Adding finalizers to rook-ceph-mon secrets
and rook-ceph-mon-endpoints configmap
We don't want to delete this resources during disaster
because these details are needed during disaster recovery

Closes: https://github.com/rook/rook/issues/8369
Signed-off-by: parth-gr <paarora@redhat.com>
2021-10-07 19:44:17 +00:00
Sébastien Han a649f64100 ceph: add signal handling for log collector
The log collector was not responding to SIGINT or SIGTERM correctly
since the parent bash process did not have the job control functionality
enabled. Now any signal received on bash will exit the container
immediately.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-27 10:50:17 +02:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 8556f3fe26 ceph: remove pool id from the peer
We don't need to put this information in the token.
It's not useful and not used anywhere.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-03 09:44:54 +02:00
Sébastien Han 630c2f6a8b ceph: add an rbd-mirror bootstrap token on cluster creation
They are scenarios where the mirroring information want to be shared
between clusters prior to creating pool. Because the bootstrap peer
import command needs a pool name to operate this is not suitable. So
additionally now each time the cluster is reconciled and on any new
clusters a new secret will be created that contains a boostrap peer
token. It can be exchanged with another cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-02 18:34:27 +02:00
Sébastien Han 7ef127b816 ceph: append additional info in the rbd-mirror bootstrap peer token
If the title looks familiar this is normal, this piece of code was
removed during https://github.com/rook/rook/pull/7604. Probably due to a
rebase. So I'm re-adding the code.

We know append additional information to the rbd-mirror bootstrap peer
token. It is useful for disaster recovery scenario where the other
cluster is reading the peer token and needs to know the pool_id as well
as the namespace.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-30 15:40:35 +02:00