Commit Graph
51 Commits
Author SHA1 Message Date
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
Sébastien Han 204ffd4065 Merge pull request #6250 from leseb/mirroring-config
ceph: add rbd-mirror configuration
2020-09-18 09:16:08 +02:00
Sébastien Han 451622a955 ceph: add rbd-mirror configuration
Rook is now capable of configuring mirroring between sites. The
implementation works at different levels:

* CephBlockPool: which introduces a new `mirroring` configuration as well
as `statusCheck`. When turned on, Rook will enable mirroring on the
pool. It will also create a bootstrap peer token and store it in a
Kubernetes Secret. The name of that Secret can be found in the Status
field of the CephBlockPool CRD. This token can be fetched and used by
other clusters to configure the site as a peer. Mirroring can be
configured either at the pool or the image level.

* CephRBDMirror: which introduces a new `peers` configuration allowing
Rook to connect to peers by passing a Secret name. The administrator will
create a Kubernetes Secret with 2 keys: 'token' for the bootstrap peer
token and 'pool' for the name of pool. Once detected the rbd-mirror
controller will go ahead and import the peer configuration.

Pool mirroring status example:

```
status:
  info:
    rbdMirrorBootstrapPeerSecretName: pool-peer-token-test
  mirroringInfo:
    lastChanged: "2020-09-17T14:47:27Z"
    lastChecked: "2020-09-17T14:48:27Z"
    summary:
      summary:
        mode: image
        peers:
        - client_name: client.rbd-mirror-peer
          direction: rx-tx
          mirror_uuid: ""
          site_name: rhcs
          uuid: c50522a4-28a4-4bd3-ba68-e11780308882
        site_name: 91eae0dd-06b1-4d2c-91f3-1311c9df382b-rook-ceph
  mirroringStatus:
    lastChecked: "2020-09-17T14:48:27Z"
    summary:
      summary:
        daemon_health: OK
        health: OK
        image_health: OK
        states:
          replaying: 1
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-17 19:05:04 +02:00
subhamkrai 25c116a4bd ceph: handle golangci-lint linter gosimple
this commit handles all the errors  check for
golangci-lint linter gosimple.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 15:10:27 +05:30
Sébastien Han bd90f9afe9 ceph: add more error handling and debug logging
The predicate was lacking from very useful errors as well as logging
messages.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
Sébastien Han e7043edbd8 ceph: reconcile if cm is config override
Earlier, we were returning only when the object was not the config
override configmap, we want the opposite.
Also, this was blocking all subsequent conditions and add the correct
check to avoid cm to reconcile.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
Travis Nielsen ce92249725 ceph: continue with memory limits below min settings
The operator will now allow the resource limits to be applied below the
recommended minimums. In small clusters, even the recommended minimums
may not be necessary. A warning is still printed to the operator log,
but we allow the configuration to continue.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-19 11:07:16 -06:00
Sébastien Han af1e8c320a ceph: do not log an error if no clusters
If the list of cluster is empty there is no need to report an error in
the logs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-08-05 10:05:51 +02:00
Travis Nielsen 06c21954e1 ceph: labels are immutable for match selectors
The labels cannot be updated on a deployment match selector.
In v1.4 a new label was added for ceph_daemon_type that is intended
to be on the pod labels, but cannot be applied to the match selectors.
Therefore, we suppress any new labels that are added to the daemons.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-04 16:47:00 -06:00
Blaine Gardner dcf8ae2526 ceph: rename pod labels function for more clarity
Rename PodLabels function to CephDaemonAppLabels for more clarity about
what the function's purpose is.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-30 11:16:30 -06:00
Blaine Gardner 9b6114c192 ceph: add 'ceph_daemon_type' label to pod labels
This can help users identify the daemon type similarly to how
'ceph_daemon_id' helps identify the daemon ID. This also makes it so
scripts users create don't have to parse "app=rook-ceph-<daemonType>"
into "<daemonType>" if they desire that bit of info.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-30 11:16:22 -06:00
subhamkrai acc4ed5df9 ceph: handling gosec errors that are not checked
this commit handles all the gosec g104
i.e audit errors not checked inside pkg.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-24 20:02:06 +05:30
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Travis Nielsen 4681f9e73d ceph: consolidate ceph config and client packages
The ceph config and client packages are conceptually the same.
To avoid circular dependencies in some cases, we simplify by combining
the packages into the client package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:43 -06:00
Sébastien Han 56954d5cc8 ceph: fix CRs not reconciling on updates
Most of the CRs expect CephCluster and CephBlockPool were not
reconciling on updates. The predicate must return true on CR diff.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-15 11:15:33 +02:00
Travis Nielsen 8f9df75f7b ceph: capture logging from the osd daemon
The OSD daemon has been missing critical flags for logging to stderr
where k8s can capture the logs. Without the --log-to-stderr=true,
all the OSD logging was essentially lost until now.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-06 14:38:37 -06:00
Sébastien Han aca541803b Merge pull request #5664 from leseb/ceph-object-user-external
ceph: add external support for object store user
2020-06-18 20:10:05 +02:00
Sébastien Han 84d1e28c99 ceph: add external support for objectstoreuser
Now, the object store user is capable of creating s3 users on an
external Ceph cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-18 16:34:01 +02:00
Sébastien Han 1e4a4b477a Merge pull request #5111 from leseb/lang-neutral-ceph
ceph: use octopus base image
2020-06-18 10:46:54 +02:00
Sébastien Han 68e62836c5 ceph: use newer octopus time format
radosgw-admin as of Octopus uses a different time format, it uses
"2006-01-02T15:04:05.999999999Z". It's close from RFC3339 but not quite
the same. This change is needed in order for the operator image (with a
Ceph Octopus based image) to perform radosgw-admin call correctly.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-17 23:00:17 +02:00
Sébastien Han c7f255a0e8 ceph: add the ability to set any pool property
We can now explicitly set any property on a given pool by using the new
Property field in the CephBlockPool Spec.

Also, this fixes the case where both `CephBlockPool` and `CephCluster`
are created at the same time. When Rook creates the pool, the cluster is
still being bootstrapped and the global option
`osd_pool_default_pg_autoscale_mode` has not bee set yet. So the pool
gets created but its `pg_autoscale_mode` property is set to `warn`
instead of `on`.

Closes: https://github.com/rook/rook/issues/5608V
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-17 10:08:00 +02:00
Ali Maredia fc579f4520 ceph: initial commit for ceph rgw multisite resources
This commit contains CR implementations for:
CephObjectRealm
CephObjectZoneGroup
CephObjectZone

Also there are changes made to the objectstore
to add rgws in the object-store to zones and
zone groups in a multisite configuration and
the removal of the --default parameter for any
realms/zonegroups/zones that are created.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-06-09 16:28:53 -04:00
Sébastien Han c8de7009de ceph: increase liveness probe start delay on OSD
When the cluster is loaded and we restart an OSD, it will need some time
to respond to socket calls, basically more to be ready.
Increasing the initialDelaySeconds of the liveness probe fixes that
issue. For OSD, it waits for 45 sec, where other daemons 10 sec.

Closes: https://github.com/rook/rook/issues/5492
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-05 15:53:49 +02:00
Travis Nielsen d1da12c6ac ceph: wait indefinitely for cleanup before removing cluster finalizer
During cluster deletion, we currently only retry for a couple minutes
to wait for the pvcs to be deleted. After the timeout, we proceed
with the cluster deletion. To properly protect the pvcs for proper
cleanup, the finalizer should not be removed until the pvcs
are all confirmed to be deleted. In order to not block other cluster
events, we re-queue the deletion event to run again every 10s
until the pvcs are deleted or the finalizer is manually removed.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-14 16:52:24 -06:00
Travis Nielsen f47bb945c2 ceph: remove duplicate controller wait setting
The controller setting for requeuing an event moved to the
opcontroller package and was no longer needed in the main
controller package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-14 16:52:24 -06:00
Madhu Rajanna 81688398f2 cleanup: use err.Wrap when the formatting is not required
Replaced err.Wrapf with err.Wrap when the formatting
is not required.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-04-29 17:42:59 +05:30
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Travis Nielsen 8d72e788c1 ceph: lookup of the cluster cannot rely on namespace name
The name of a CephCluster is commonly the same as the namespace,
but not always. When looking up the ceph cluster we can look for
the first one in the namespace rather than requiring the name
to be the same as the namespace.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-22 13:42:40 -06:00
Sébastien Han c1b8a7aa9e ceph: extract rbd-mirror to its own crd
Previously, the rbd-mirror daemon was integrated into the `CephCluster`
CRD. This wasn't really practical since we would have to wait for the
whole orchestration to be done to actually set it up. The same goes for
any CR update. Let's say you want to change the number of daemons, Rook
would go through mons, mgrs and osds until it get to rbd-mirror.
This triggers an undesired full orchestration.

With its own CRD this component just gains a lot more flexibility.

Closes: https://github.com/rook/rook/issues/5084
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-15 09:12:56 +02:00
Travis Nielsen ff1ab64d8d ceph: skip reconcile events where the spec not updated
The controllers should ignore events for status updates. The reconcile
only needs to happen when the spec is updated or the resource is marked
for deletion. Otherwise, the controllers may stay in an update loop as
the reconcile events are triggered again every time there is a status
update during the reconcile.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-09 11:20:45 -06:00
Sébastien Han 71799682ca ceph: print the reason for reconciling
General debugging helper without having to turn on DEBUG logs.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-06 16:07:01 +02:00
Travis Nielsen eeb687ad14 ceph: log skipped reconcile based on ceph health
When the Ceph health is HEALTH_ERR the controllers will skip the
reconcile. With info level logging there needs to be an indication
of this decision to skip the reconcile, otherwise users will
wonder why their resources aren't being created.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-06 16:07:01 +02:00
Sébastien Han c5075d55d1 ceph: trigger update on child controllers
When the CephCluster CR gets updated with a new image version, we need
to notify the child controllers of that upgrade.
For this, we set a 'ceph_version' label on the CR itself which will
effectively trigger a reconcile based on an update event.
This needs to be revisited and see how we can make use of
`EnqueueRequestsFromMapFunc` which might be a better approach using
controller-runtime built-in.

Closes: https://github.com/rook/rook/issues/5153
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-03 13:56:49 -06:00
Travis Nielsen 96dd92cc7f ceph: change verbose info log to debug in controller
Keeping the operator log clean...

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-04-03 09:15:25 -06:00
Sébastien Han ed37a6a699 Merge pull request #5128 from leseb/liveness-other-daemons
ceph: add liveness probe to mon, mds and osd daemons
2020-04-01 10:53:12 +02:00
Sébastien Han 040193bb5a ceph: remove DaemonType type
This type was a string already and was just making us doing string()
calls all the time to it's not worth it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-01 09:08:18 +02:00
Sébastien Han 7788901a88 ceph: add liveness probe to mon, mds and osd daemons
Now Kubernetes will perform liveness checks on mon, mds and osd daemons.
The command will:

* call the socket (check for existence)
* execute a command and check the return code (success if 0)

This handles the case where the daemon is stuck locally and
unresponsive. It's unlikely but not impossible.
These checks bring more robustness to the implementation.

rbd-mirror and nfs have been leftover for the following reason. The
rbd-mirror socket name is different from other daemons (could be fixed
though): /run/ceph/ceph-client.rbd-mirror.a.1.94362516231272.asok also,
the command to call would need to be changed from "status" to "rbd
mirror status" so we can keep this for a later.
The nfs ganesha has no socket only a PID file which doesn't mean much.
No PID means the process does not run so Kubernetes will already handle
this and the pod will crash loop.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-31 15:20:42 +02:00
Travis Nielsen 0e55aa348b tests: reduce unhelpful debug logging
Removed some extremely verbose debug logging that makes
the logs difficult to read when debug is enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-30 13:32:47 -06:00
Sébastien Han 8db14885b5 ceph: controller fix misleading debug log
They are cases where we let the exponential backoff retry and some where
we force our own retry. Let's properly log this information.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 19:12:17 +01:00
Sébastien Han c84d66de1b ceph: controller, reconcile faster
Let's not wait for the CephCluster to be done reconciling but instead
check for the Ceph cluster status, if it's closed to "ok" then we
proceed so HEALTH_OK and HEALTH_WARN are accepted.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 18:45:19 +01:00
Sébastien Han b5e50481ff ceph: controller: retry after 10sec even if no cluster
Even if there is no CephCluster we still want to retry every 10sec as
one will likely show up soon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Sébastien Han 7b544f04d4 ceph: reconcile more often when cluster is not ready
Let's not apply the exponential backoff when waiting for the cluster to
be ready, let's only apply it when there is no cluster.

Closes: https://github.com/rook/rook/issues/5059
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-25 17:24:30 +01:00
Sébastien Han 0e2f84f998 ceph: convert NFS controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4941
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-20 17:38:03 +01:00
Sébastien Han ca0a30f38d ceph: convert Filesystem controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-19 23:34:58 +01:00
Travis Nielsen 18b0e7d295 ceph: scrub ceph commands to write actions to the log
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Sébastien Han f268c897e9 ceph: Convert the Ceph ObjectStore controller to the controller-runtime
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.

Closes: https://github.com/rook/rook/issues/4937
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-17 15:12:41 -06:00
Umanga Chapagain 06ba93e629 Ceph: refactor GetOperatorSetting for reuse
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-09 21:41:56 +05:30
Sébastien Han 9f2867e12a ceph: separate controller for CephObjectStoreUser CRD
Now, the CephObjectStoreUser CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-06 11:53:40 +01:00
Sébastien Han a137b31e1a ceph: refactor controller helper
Clean and refactor code helper for controller-runtime.
Implement those into the block pool controller.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-05 12:37:44 -07:00