Commit Graph
96 Commits
Author SHA1 Message Date
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
Travis Nielsen 95b84a6c12 Merge pull request #11099 from avanthakkar/fix-configure-rbd-stat-pools
operator: Don't remove existing pools for mgr/prometheus/rbd_stats_pools
2022-10-12 14:20:47 -06:00
Avan Thakkar e8b74207dd operator: don't remove existing pools for mgr/prometheus/rbd_stats_pools
Rook currently knows rbd pools only if they defined through CephBlockPool, so it overrides the config
mgr/prometheus/rbd_stats_pools which may have already been set for pools external to Rook. Check config
before running set command and if there some existing pools append them to enableStatsForCephBlockPools.

Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2022-10-12 20:59:44 +05:30
Travis Nielsen 4c11ba76d7 Merge pull request #10721 from zhucan/bugfix-10712
pool: add timeout to ceph cmd
2022-10-12 08:31:52 -06:00
parth-gr 26584fc6e5 core: update loadclusterInfo with multus check
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor

Signed-off-by: parth-gr <paarora@redhat.com>
2022-09-22 14:46:51 +05:30
zhucan aa41742c83 pool: add timeout to rbd cmd
Signed-off-by: zhucan <zhucan.k8s@gmail.com>
2022-09-16 10:26:20 +08:00
Rakshith R c02f3194d1 pool: initialize only rbd application pools
`rbd pool init` cmd initializes pool for rbd images.
This commit makes modification to init only
rbd application pools, since it is not required
by other pools like ".nfs",".mgr" & "mgr_devicehealth".

Signed-off-by: Rakshith R <rar@redhat.com>
2022-09-13 18:18:26 +05:30
Travis Nielsen 89fadabcc1 Merge pull request #10575 from shalevpenker97/EnableRBDStats-mgr-fix
Fix EnableRBDStats mgr
2022-07-18 12:57:54 -06:00
shalev.p 49996522bf pool: configureRBDStats remove postfix dot in mgr config path
The ceph rbd_stats does not take affect with dot postfix.

Closes: https://github.com/rook/rook/issues/10574
Signed-off-by: shalev.p <shalev.p@taboola.com>
2022-07-11 18:21:42 +03:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Travis Nielsen e3ca5bfbcc pool: delete ceph pool when blockpool cr is deleted
During deletion of a CephBlockPool CR, the underlying ceph
pool was not being deleted. During the deletion sequence
this is due to the spec being refreshed upon a call to
check for dependents in ReportDeletionNotBlockedDueToDependents().
The pool name was then coming up blank during the deletion
request and rook of course then wasn't finding the pool
to delete.

Now Rook properly sets the pool name in the named spec
every time it is converted internally from a CephBlockPool
type to a NamedPoolSpec.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-06-01 10:20:57 -06:00
Sébastien Han 583791c45c core: move clusterInfo code to the controller package
The CSI package needs to load clusterInfo, today this code is in the mon
package which makes the call of LoadClusterInfo impossible without
having a circular import.

Signed-off-by: Sébastien Han <seb@redhat.com>
2022-04-26 11:05:02 +02:00
Madhu Rajanna b0fc7c9b92 namespace: add new CRD
This introduces a new CRD to add the ability
to create rados namespace for a given
ceph block pool. Typically the name of the pool
is the name of the blockpool created by rook.

Closes: #7035

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2022-04-05 10:10:04 +05:30
Divyansh Kamboj 9008409f87 core: add context parameter to functions
This commit adds context parameter to various functions, and remove the
usage of context.TODO.

Closes: https://github.com/rook/rook/issues/8701
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
2022-03-22 08:19:07 +05:30
Sébastien Han 18b19de639 Merge pull request #9687 from parth-gr/observedGeneration
core: add observedGeneration to CR status
2022-03-17 10:36:11 +01:00
Travis Nielsen ad99f9e205 pool: update the quincy default pool to .mgr
In Quincy, the device_health_metrics pool is renamed to .mgr.
In the example cluster-test.yaml, we therefore update the
example to create the pool with the new name.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-03-16 14:18:49 -06:00
parth-gr 2dfd64a97c core: add observedGeneration to CR status
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller

Closes: https://github.com/rook/rook/issues/9673

Signed-off-by: parth-gr <paarora@redhat.com>
2022-03-16 19:58:35 +05:30
subhamkrai 6a9ff434d2 core: add k8s events in controller reconciler
adding k8s event in controller reconciler when
1. when reconciler starts
2. when deletion reconcile triggered.

Closes: https://github.com/rook/rook/issues/9462
Signed-off-by: subhamkrai <srai@redhat.com>
2022-03-14 18:55:26 +05:30
Travis Nielsen 221b3dfb72 object: update pool properties during reconcile
The reconcile was skipping updating most pool properties for
object stores. The implementation of pools between the file,
object, and pool controllers had some duplicate code, so
this change also factors out the common code for better
reuse in a single place. Anytime a pool is created or updated,
it will now consistently update all the pool properties
that are expected to be modifiable.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-02-04 13:41:12 -07:00
Travis Nielsen 7ee9cc9d56 pool: allow configuration of built-in pools with non-k8s names
The built-in pools device_health_metrics and .nfs created by ceph
need to be configured for replicas, failure domain, etc.
To support this, we allow the pool to be created as a CR.
Since K8s does not support underscores in the resource names
the operator must translate this special pool name into
the name expected by ceph.

This also sets the basis for allowing filesystem data
pools to specify the desired pool name instead of requiring
a generated name.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-15 15:32:02 -07:00
Blaine Gardner 0c76be8c38 pool: file: object: clean up health checkers for force deletion
When CephBlockPool, CephFilesystem, or CephObjectStore resources are
deleted after removing their finalizer, the code path to stop monitoring
was not stopping monitoring since a non-present resource does not have a
name and namespace attached. When the object is deleted, ensure the
internal representation used to stop monitoring has a name and namespace
to fix the issue.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-13 16:07:01 -07:00
Blaine Gardner 03ba7dec64 pool: file: object: clean up stop health checkers
Clean up the code used to stop health checkers for all controllers
(pool, file, object). Health checkers should now be stopped when
removing the finalizer for a forced deletion when the CephCluster does
not exist. This prevents leaking a running health checker for a resource
that is going to be imminently removed.

Also tidy the health checker stopping code so that it is similar for all
3 controllers. Of note, the object controller now uses namespace and
name for the object health checker, which would create a problem for
users who create a CephObjectStore with the same name in different
namespaces.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-19 10:29:12 -07:00
Yuichiro Ueno 0b575703c7 core: add context parameter to k8sutil deployment
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 14:58:13 +09:00
Travis Nielsen fd10d98dc6 core: treat cluster as not existing if the cleanup policy is set
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-27 10:25:06 -06:00
Blaine Gardner 2f850b6ae6 Merge pull request #8613 from subhamkrai/remove-nautilus
ceph: remove ceph nautilus, ceph octopus to default
2021-10-25 09:09:18 -06:00
Yuichiro Ueno 3fd86f83ae core: add context parameter to opcontroller
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-10-25 20:45:06 +09:00
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
Rakshith R ab87e1d238 ceph: initialize rbd block pool after creation
This is done in order to prevent deadlock when parallel
PVC create requests are issued on a new uninitialized
rbd block pool due to https://tracker.ceph.com/issues/52537.

Fixes: #8696

Signed-off-by: Rakshith R <rar@redhat.com>
2021-10-06 19:54:51 +05:30
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 4cb9b53487 ceph: fix pool deletion when running on multus
The ceph-block-pool controller was missing the network spec from the
cephcluster that contains all the details about networking.
So when the pool was deleting on multus the controller was not proxying
the rbd command to the mgr pod but executed the command in the operator
pod which does not have the network annotation and then connect to the
ceph cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-08 16:10:31 +02:00
Sébastien Han d4dd0577f9 Merge pull request #8529 from sp98/id-mapping
ceph: add ClusterID and PoolID mappings between local and peer cluster
2021-08-31 17:32:14 +02:00
Santosh Pillai 3f8abec403 ceph: add ClusterID and PoolID mappings between local and peer cluster
During disaster recovery/migration of a cluster, as part of the failover, the
kubernetes artifacts like deployment, PVC, PV, etc will be restored to a new
cluster by the admin. Even if the kubernetes objects are restored the
corresponding RBD/CephFS subvolume cannot be retrieved during CSI operations as
the clusterID and poolID are not the same in both clusters

This PR creates a mapping between Cluster ID and RBD Pool ID between
local cluster and peer cluster.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-08-31 20:08:56 +05:30
Travis Nielsen cdfe1982d1 ceph: consolidate the calls to set mon config
The mon config had two different implementations that have evolved
over the lifetime of the project. This is a simple refactor to remove
the SetConfig() option and stick with the MonStore as a single
implementation for updating the mon store.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-08-24 14:38:26 -06:00
Sébastien Han 2d55e69416 ceph: move scheme initialization to the same place
Let's initialize the schemes in a single place instead of doing it
when each controller initializes.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-06 11:33:10 +02:00
Sébastien Han 630c2f6a8b ceph: add an rbd-mirror bootstrap token on cluster creation
They are scenarios where the mirroring information want to be shared
between clusters prior to creating pool. Because the bootstrap peer
import command needs a pool name to operate this is not suitable. So
additionally now each time the cluster is reconciled and on any new
clusters a new secret will be created that contains a boostrap peer
token. It can be exchanged with another cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-02 18:34:27 +02:00
Santosh Pillai 3b50f5fd00 ceph: add MirroringPeerSpec to CephBlockPool.Spec.Mirroring
Instead of having the RBDMirroringPeerSpec in the RBDMirror CRD we want
to move it on the CephBlockPool CRD so that each pool can have its own peer.
This enables pool to re-use an existing peer secret if it points to the
same cluster peer.So we don't need to duplicate secret peer anymore.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-07-28 17:47:36 +05:30
Sébastien Han baaea4a1ea Merge pull request #7604 from leseb/cephfs-mirror-peer-config
ceph: add filesystem mirror peers configuration
2021-07-05 10:36:30 +02:00
Sébastien Han d59ae25d2b Merge pull request #8215 from sp98/disable-mirroring
ceph: ability to disable pool mirroring
2021-07-05 10:22:24 +02:00
Santosh Pillai 05b0ae09bb ceph: disable mirroring
The PR disables mirroring on a pool when PoolSpec.Mirroring.Enabled
is set to false
- disable mirroring on the pool for `Mirroring.Mode == pool`
- Add warning for the user to disable mirroring manually for `Mirroring.Mode == image`
- Stop the mirroring health checker goroutine
- Reset the mirroring health check status in the CR.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-07-02 10:35:38 +05:30
Sébastien Han b578f916e7 ceph: add fs mirror config
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.

So the automatic configuration of Ceph Filesystem peers is now possible.

By editing the CephFilesystem CRD, you can now turn on mirroring:

```yaml
  mirroring:
    enabled: false
    # list of Kubernetes Secrets containing the peer token
    # for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
    peers:
      secretNames:
        - secondary-cluster-peer
```

Also, the mirroring status is displayed in the CR status:

```
status:
  info:
    fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
  mirroringStatus:
    daemonsStatus:
    - daemon_id: 4186
      filesystems:
      - filesystem_id: 2
        name: myfs
    lastChecked: "2021-07-01T14:16:29Z"
  phase: Ready
  snapshotScheduleStatus:
    lastChecked: "2021-07-01T14:16:29Z"
    snapshotSchedules:
    - fs: myfs
      path: /
      rel_path: /
      retention: {}
      schedule: 24h
```

Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 17:35:19 +02:00
Sébastien Han 8c589a3785 ceph: silence harmless errors
Let's not look at misleading errors if the operator is still
initializing.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 15:37:08 +02:00
Santosh Pillai 8d48a9e10a ceph: update blockPoolChannel before starting the mirror monitoring
update blockPoolChannel monitoringRunning field to `true` before
starting the goroutine to monitor mirroring

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2021-07-01 15:27:37 +05:30
Travis Nielsen b1f4411fb4 ceph: disable insecure global id if no insecure clients
In the latest Ceph releases starting with v16.2.1, all clients are recommended
to be updated so they will have a security fix to connect with a secure
global ID. A health warning will be raised if any insecure clients are connected
and another health warning is raised if insecure clients are still allowed.
Rook will now disable allowing the insecure clients if the health warning
is not being raised to indicate that there are insecure clients still connected.
This means that upgraded clusters will not have this disabled until all the
daemons are updated.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-06-10 10:05:30 -06:00
Sébastien Han 90bea8a560 ceph: stop using radosgw-admin CLI for s3 user management
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.

Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 11:08:22 +02:00
Sébastien Han e478bdadb1 ceph: stop the monitoring before cleanup
Saw this today in the logs:

```
2021-06-02 13:02:54.787838 I | ceph-block-pool-controller: deleting pool "testpool"
2021-06-02 13:02:56.178094 I | cephclient: no images/snapshosts present in pool "testpool"
2021-06-02 13:02:56.178125 I | cephclient: purging pool "testpool" (id=17)
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x0 pc=0x1cef49f]

goroutine 1607 [running]:
github.com/rook/rook/pkg/operator/ceph/pool.(*mirrorChecker).checkMirroringHealth(0xc001a98ba0, 0x8, 0xc000d34c60)
	/home/runner/work/rook/rook/pkg/operator/ceph/pool/health.go:115 +0xdf
github.com/rook/rook/pkg/operator/ceph/pool.(*mirrorChecker).checkMirroring(0xc001a98ba0, 0xc0016fd500)
	/home/runner/work/rook/rook/pkg/operator/ceph/pool/health.go:82 +0x147
created by github.com/rook/rook/pkg/operator/ceph/pool.(*ReconcileCephBlockPool).reconcile
	/home/runner/work/rook/rook/pkg/operator/ceph/pool/controller.go:298 +0xee5
```

Essentially, it's intermittent but when deleting the pool the
healthcheck kicked in and fetched the mirroring status, which returned
empty. THe subsequent code tried to access content of a nil pointer,
  hence the error.
So now we stop monitoring first, then we proceed with the deletion.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-02 16:14:10 +02:00
Sébastien Han dc0985c7f6 ceph: rehydrate the bootstrap peer token secret on monitor changes
If the monitors addresses change, we must update the peer secret with
the new addresses. For this, each time the configmap
"rook-ceph-mon-endpoints" is updated, we trigger a reconcile on the
CephBlockPool CRs, which will update the token secret.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-20 08:50:22 +02:00
Sébastien Han 2793750146 ceph: silence harmless errors when the operator is not initialized
If the ceph cli outputs an error with "error calling conf_read_file"
this means that the operator has not written its ceph configuration
file. Thus ceph cli commands will fail, so we can just ignore that since
the operator will soon write this file in its initialization sequence.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-18 19:35:07 +02:00
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
Sébastien Han 97be23e374 ceph: apply finalizer before updating object status
We must add the finalizer right after the object creation otherwise the
seerver will later return an error on update that the object has been
modified. Indeed, it has been by the task that updates the status when
the object is first created.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-18 18:16:59 +01:00
Sébastien Han afc7ecff31 ceph: add snapshot scheduling for mirrored pools
Now, we can schedule snapshots on pools from the CephBlockPool CR when
the pool is mirrored.
It can be enabled like this:

```
mirroring:
  enabled: true
  mode: pool
  snapshotSchedules:
    - interval: 24h # daily snapshots
      startTime: 14:00:00-05:00
```

Multiple schedules are supported since snapshotSchedules is a list.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-18 09:50:09 +01:00