Commit Graph
113 Commits
Author SHA1 Message Date
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Omar Pakker 8f9055809f osd: add privileged support (back) to blkdevmapper securityContext (work-around)
The blockdevmapper securityContext was changed to request a minimal set of
required capabilities for its operation and drop running as privileged.
While the base change works and is valid in terms of the container's copy operation,
it turns out that OpenShift may require some additional configuration not
currently covered by the limited securityContext and the capabilities granted.

To not break those OpenShift deployments, make the blkdevmapper securityContext
listen to the ROOK_HOSTPATH_REQUIRES_PRIVILEGED flag again to set privileged mode.
This flag is true on OpenShift deployments and running as privileged
works around the (missing) configuration problem for now.
To properly drop privileged completely some additional investigation needs
to be done on OpenShift deployments without relying on privileged execution.

Signed-off-by: Omar Pakker <Omar007@users.noreply.github.com>
2021-11-17 12:25:13 +01:00
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
Yuval Lifshitz 71ed45b69b rgw: implement bucket notifications for object storage
following the design from here:
https://github.com/rook/rook/blob/master/design/ceph/object/ceph-bucket-notification-crd.md

Closes: https://github.com/rook/rook/issues/5313
Signed-off-by: Yuval Lifshitz <ylifshit@redhat.com>
2021-11-04 11:20:40 +02:00
Travis Nielsen fd10d98dc6 core: treat cluster as not existing if the cleanup policy is set
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-27 10:25:06 -06:00
Yuichiro Ueno 3fd86f83ae core: add context parameter to opcontroller
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-10-25 20:45:06 +09:00
parth-gr 7c99858a77 ceph: add finalizers to rook-ceph-mon secrets and configmap
Adding finalizers to rook-ceph-mon secrets
and rook-ceph-mon-endpoints configmap
We don't want to delete this resources during disaster
because these details are needed during disaster recovery

Closes: https://github.com/rook/rook/issues/8369
Signed-off-by: parth-gr <paarora@redhat.com>
2021-10-07 19:44:17 +00:00
Sébastien Han a649f64100 ceph: add signal handling for log collector
The log collector was not responding to SIGINT or SIGTERM correctly
since the parent bash process did not have the job control functionality
enabled. Now any signal received on bash will exit the container
immediately.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-27 10:50:17 +02:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 8556f3fe26 ceph: remove pool id from the peer
We don't need to put this information in the token.
It's not useful and not used anywhere.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-03 09:44:54 +02:00
Sébastien Han 630c2f6a8b ceph: add an rbd-mirror bootstrap token on cluster creation
They are scenarios where the mirroring information want to be shared
between clusters prior to creating pool. Because the bootstrap peer
import command needs a pool name to operate this is not suitable. So
additionally now each time the cluster is reconciled and on any new
clusters a new secret will be created that contains a boostrap peer
token. It can be exchanged with another cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-02 18:34:27 +02:00
Sébastien Han 7ef127b816 ceph: append additional info in the rbd-mirror bootstrap peer token
If the title looks familiar this is normal, this piece of code was
removed during https://github.com/rook/rook/pull/7604. Probably due to a
rebase. So I'm re-adding the code.

We know append additional information to the rbd-mirror bootstrap peer
token. It is useful for disaster recovery scenario where the other
cluster is reading the peer token and needs to know the pool_id as well
as the namespace.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-30 15:40:35 +02:00
Satoru Takeuchi 30e4fbb01f ceph: make the timeout of ceph commands cofigurable
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-27 12:42:16 +00:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Sébastien Han 1e45eaa436 ceph: always rehydrate the access and secret keys
Prior to that the access and secret keys were left empty if the user
already existed, which led to updating the secret with empty values when
the operator restarts.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-09 10:39:52 +02:00
Sébastien Han baaea4a1ea Merge pull request #7604 from leseb/cephfs-mirror-peer-config
ceph: add filesystem mirror peers configuration
2021-07-05 10:36:30 +02:00
Travis Nielsen c39c1c7ddf ceph: retry reconcile immediately after cancellation
If the reconcile is cancelled due to a CR update, we want to retry the next
reconcile immediately rather than wait for the exponential backoff timeout
if the reconcile was already failing. The wait can easily be minutes
if the reconcile was in this state, which makes it appear the operator
is ignoring the request to start a new reconcile.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-02 08:23:25 -06:00
Sébastien Han b578f916e7 ceph: add fs mirror config
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.

So the automatic configuration of Ceph Filesystem peers is now possible.

By editing the CephFilesystem CRD, you can now turn on mirroring:

```yaml
  mirroring:
    enabled: false
    # list of Kubernetes Secrets containing the peer token
    # for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
    peers:
      secretNames:
        - secondary-cluster-peer
```

Also, the mirroring status is displayed in the CR status:

```
status:
  info:
    fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
  mirroringStatus:
    daemonsStatus:
    - daemon_id: 4186
      filesystems:
      - filesystem_id: 2
        name: myfs
    lastChecked: "2021-07-01T14:16:29Z"
  phase: Ready
  snapshotScheduleStatus:
    lastChecked: "2021-07-01T14:16:29Z"
    snapshotSchedules:
    - fs: myfs
      path: /
      rel_path: /
      retention: {}
      schedule: 24h
```

Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 17:35:19 +02:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Sébastien Han 153f1d661c ceph: append additional info in the rbd-mirror bootstrap peer token
We know append additional information to the rbd-mirror bootstrap peer
token. It is useful for disaster recovery scenario where the other
cluster is reading the peer token and needs to know the pool_id as well
as the namespace.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-23 18:49:39 +02:00
Blaine Gardner 21e290e003 ceph: implement dependencies for CephCluster
Implement the first step of `design/ceph/resource-dependencies.md` to
add dependency checking when deleting a CephCluster.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-10 10:06:10 -06:00
Sébastien Han e17fa5be51 Merge pull request #7998 from leseb/replace-radosgw-admin-cli-with-goceph
ceph: stop using radosgw-admin CLI for s3 user management
2021-06-09 16:01:14 +02:00
Sébastien Han 90bea8a560 ceph: stop using radosgw-admin CLI for s3 user management
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.

Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 11:08:22 +02:00
Blaine Gardner 386eeb7eb9 ceph: fix detection of delete event reconciliation
Fix a bug where delete events on Rook-Ceph CRs would be detected
multiple times if deletion was blocked. Old method used pointer
comparison instead of using the (metav1.Time) Equal() method.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-08 16:48:53 -06:00
Sébastien Han db91fd8bc3 ceph: do not configure external metric endpoint is not present
If no endpoint are configured let's simply return.

Closes: https://github.com/rook/rook/issues/7963
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-24 14:41:26 +02:00
Sébastien Han 5e1e9f4c92 ceph: actively update the service endpoint for external mgr
If the cluster is external we want to periodically rehydrate the mgr
endpoint. This handles the scenarion where the active manager changes,
 so we need to update the endpoint with the new IP address.
The create-external-cluster-resources.py script now requires an extra
permission to query the manager service so Rook can discover the active
one and its IP.

Testing:

```
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
    "epoch": 37,
    "available": true,
    "active_name": "b",
    "num_standby": 1
}

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME                     ENDPOINTS          AGE
rook-ceph-mgr-external   172.17.0.12:9283   3m10s

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] k scale --replicas=0 deployment rook-ceph-mgr-b
deployment.apps/rook-ceph-mgr-b scaled

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
    "epoch": 40,
    "available": true,
    "active_name": "a",
    "num_standby": 0
}

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME                     ENDPOINTS          AGE
rook-ceph-mgr-external   172.17.0.13:9283   3m55s
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-11 16:08:40 +02:00
parth-gr bb23622b72 core: enhancement in debug logs
When DEBUG logging is enabled in the operator, there are a number of overwhelming messages in the log that are very overwhelming and don't seem useful, which makes debug mode difficult to use
Updated and Cutted down the messages that are not useful in debug mode

Closes: https://github.com/rook/rook/issues/7499
Signed-off-by: parth-gr <paarora@redhat.com>
2021-04-28 22:28:10 +05:30
Blaine Gardner 9d657f464c ceph: redact secret info from reconcile diffs
To avoid logging sensitive information, do not output diffs from secrets
when reconciling resources.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-04-15 08:14:34 -06:00
Lars Lehtonen 41c567beaf ceph: fix multiple imports
This fixes additional double-imports.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-13 02:02:09 -07:00
Travis Nielsen 721acd1a8a Revert "ceph: added rook-ceph-default service account"
This reverts commit 737fb099fe.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-04-08 11:46:12 -06:00
Sébastien Han 51e7310e09 ceph: add dualstack support on pacific
With Pacific comes the support for dualstack where ceph daemons can
listen on both ipv4 and ipv6 stacks.
A new field in the network spec has been added: `dualStack`

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-31 18:04:29 +02:00
Travis Nielsen 3c37def52f Merge pull request #7406 from parth-gr/service-account
ceph: added rook-ceph-default service account
2021-03-30 08:06:40 -06:00
Sébastien Han 0c799c3b8f Merge pull request #7476 from subhamkrai/ceph-versions
ceph: update cephCluster CR with ceph versions output
2021-03-30 09:10:01 +02:00
Blaine Gardner 1df4336a68 Merge pull request #7386 from BlaineEXE/update-osds-in-parallel
ceph: Update osds in parallel
2021-03-29 16:32:07 -06:00
Blaine Gardner 795124b7a8 ceph: update osds in parallel
Update OSDs in parallel per the design in
design/ceph/update-osds-in-parallel.md

The max number of OSDs updated in parallel is currently fixed at 20.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-03-29 10:55:28 -06:00
Blaine Gardner 660d2fe442 Merge pull request #7493 from BlaineEXE/codespell-tidy
test: document or clean up codespell flags
2021-03-29 10:07:53 -06:00
subhamkrai 05d4c2776c ceph: update cephCluster CR with ceph versions output
update cephCluster CR with ceph versions output.
this output will contain ceph version of the
ceph daemons which will help with upgrade status.

ceph versions command output
```
ceph versions
{
    "mon": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 3
    },
    "mgr": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "osd": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "mds": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 2
    },
    "rbd-mirror": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "rgw": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 1
    },
    "overall": {
        "ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable)": 9
    }
}
```

CR status
```
status:
  ceph:
   ---
    versions:
      mds:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 2
      mgr:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      mon:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 3
      osd:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      overall:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 9
      rbd-mirror:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
      rgw:
        ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable): 1
```

Signed-off-by: subhamkrai <srai@redhat.com>
2021-03-29 21:04:35 +05:30
parth-grandTareq Sharafy 737fb099fe ceph: added rook-ceph-default service account
When a private docker registry is used and an image pull secret is specified in the chart, the pods with default Service Account fail to pull the image due to authentication issues.
Added rook-ceph-default service account and modify the pods specifications by adding the serviceAccountName.

Closes: https://github.com/rook/rook/issues/6673
Co-authored-by: Tareq Sharafy <tareq.sha@gmail.com>
Signed-off-by: parth-gr <partharora1010@gmail.com>
2021-03-29 20:20:40 +05:30
Blaine Gardner 71c41180be test: document or clean up codespell flags
Document codespell flags that are necessary. Remove codespell flags with
minor code changes if possible.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-03-26 15:00:49 -06:00
Satoru Takeuchi 9eb3160d8b ceph: validate all owner references
Remaining work of #7259

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-26 06:43:54 +00:00
Travis Nielsen c23238cddb ceph: refactor integration tests for simplification
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 11:26:10 -06:00
Travis Nielsen 64e28af741 ceph: allow flex driver and discovery to be enabled with configmap
For testing purposes, we need to configure the flex driver
and discovery daemon with the operator settings configmap.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-17 08:39:26 -06:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Travis Nielsen d93ff428ec ceph: reconcile the mgr services with the active mgr
The active mgr should match the labels on the services that
are available for prometheus and the dashboard. The selector
labels must be updated whenever there is a new active mgr.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-10 08:29:16 -07:00
Travis Nielsen 15d0537cec ceph: only raise conditions that represent current state
The conditions on the cephcluster CR were being set to false
when not in progress, which is misleading to the purpose of
conditions. The conditions will now always represent current
state of the cluster. Transient conditions that show progress of
a reconcile will be removed from the status.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-02 14:15:47 -07:00
Sébastien Han 8d033efb5a ceph: silence harmless errors
If the ceph cli outputs an error with "error calling conf_read_file"
this means that the operator has not written its ceph configuration
file. Thus ceph cli commands will fail, so we can just ignore that since
the operator will soon write this file in its initialization sequence.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-02-03 17:28:44 +01:00
Travis Nielsen 5371c943ca ceph: simplify the log-collector container name
The log-collector container name was exceeding the 63-char limit
if the parent CR name was too long. The container name just needs
to be unique to the pod spec, so we simplify it to remove the parent
name and avoid hitting the limit in the pod spec.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-02-02 12:42:53 -07:00
Sébastien Han c0123cf182 ceph: add cephfs mirroring support
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.

The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050

The same goes for integration tests, they will get added later once we
start testing on Pacific.

Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-28 19:21:18 +01:00
Sébastien Han ebbf332d8d ceph: update rgw and mds deployment for logCollector
If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.

Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-22 18:19:07 +01:00