The toolbox can now interact with S3 gateways using the `s5cmd` tool.
The binary is only 12M so this does not add up too much to the operator
image size.
Closes: https://github.com/rook/rook/issues/4968
Signed-off-by: Sébastien Han <seb@redhat.com>
The rook operator as well as the toolbox pod run with the "rook" user
with UID 2016. The UID was chosen based on the year of the initial
commit in the rook/rook repository.
No more root user running.
Closes: https://github.com/rook/rook/issues/8734
Signed-off-by: Sébastien Han <seb@redhat.com>
Rook cluster-wide encryption can now use the native Kubernetes
authentication to interact with vault KMS instead of using the token
method.
Signed-off-by: Sébastien Han <seb@redhat.com>
Adding finalizers to rook-ceph-mon secrets
and rook-ceph-mon-endpoints configmap
We don't want to delete this resources during disaster
because these details are needed during disaster recovery
Closes: https://github.com/rook/rook/issues/8369
Signed-off-by: parth-gr <paarora@redhat.com>
We don't need to use tini.
We don't have anything in the rook operator that would
either create zombie processes (no threads) or use
exec (to fork). The Go binary has a really good
signal handling mechanism.
Closes: https://github.com/rook/rook/issues/8794
Signed-off-by: Sébastien Han <seb@redhat.com>
In Rook v1.8 the min version of K8s supported is updated to 1.16.
Users running on older versions of K8s are recommended to update
to 1.16 or newer before updating to Rook v1.8.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This PR updates the required RBAC, templates,
CSI image version and examples for new
cephcsi v3.4.0 release.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Instead of having the RBDMirroringPeerSpec in the RBDMirror CRD we want
to move it on the CephBlockPool CRD so that each pool can have its own peer.
This enables pool to re-use an existing peer secret if it points to the
same cluster peer.So we don't need to duplicate secret peer anymore.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Multiple filesystems are supported as of the Ceph Pacific release. So we
remove the flag from the operator configuration and just have a ceph
version check instead.
The Operator configuration option `ROOK_ALLOW_MULTIPLE_FILESYSTEMS`
has been removed in favor of simply verifying the Ceph version is
at least Pacific.
Multiple filesystems are stable since Ceph Pacific.
So users who had `ROOK_ALLOW_MULTIPLE_FILESYSTEMS` enabled will
need to update their Ceph version to Pacific.
Closes: https://github.com/rook/rook/issues/7183
Signed-off-by: Sébastien Han <seb@redhat.com>
Recently, the builds of `ceph/ceph` image moved to quay.io, see
https://github.com/ceph/ceph-build/pull/1883 for more details.
Current images will remain but new builds will happen on quay.io only.
This means that tags such as `v14.2`, `v15.2`,`v16.2` will need to
switch to quay.io to get updates.
Signed-off-by: Sébastien Han <seb@redhat.com>
The CRDs v1 requires the full schema for all settings, so we now
generate the CRDs for nfs for full fidelity of all settings.
Co-authored-by: Nicolaj Græsholt <figaw@hotmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRDs v1 requires the full schema for all settings, so we now
generate the CRDs for cassandra for full fidelity of all the
settings.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The mon daemon in a stretch cluster now can have its location set
as a CLI param instead of setting it with a separate command.
This enables mon failover to set the location of a mon immediately
when it is joining quorum instead of having a delayed command
to set the location.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.
So the automatic configuration of Ceph Filesystem peers is now possible.
By editing the CephFilesystem CRD, you can now turn on mirroring:
```yaml
mirroring:
enabled: false
# list of Kubernetes Secrets containing the peer token
# for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
peers:
secretNames:
- secondary-cluster-peer
```
Also, the mirroring status is displayed in the CR status:
```
status:
info:
fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
mirroringStatus:
daemonsStatus:
- daemon_id: 4186
filesystems:
- filesystem_id: 2
name: myfs
lastChecked: "2021-07-01T14:16:29Z"
phase: Ready
snapshotScheduleStatus:
lastChecked: "2021-07-01T14:16:29Z"
snapshotSchedules:
- fs: myfs
path: /
rel_path: /
retention: {}
schedule: 24h
```
Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
Creates hybrid crush rule for choosing Primary OSD for
high performing SSD devices and remaining OSD for low performance HDD devices.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Implement the first step of `design/ceph/resource-dependencies.md` to
add dependency checking when deleting a CephCluster.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Since Pacific, we can deploy multiple object gateway using the same
keying and they will appear in the service map separetly. So from now on
and on upgrades to Pacific Rook will remove all extra deployments to
only keep 1. In this single deployment the number of replica will be set
the desired `instances` count from the CRD spec.
You will now see gateways like this on Pacific:
```
rgw: 3 daemons active (1 hosts, 1 zones)
```
See an upgrade operator logs from Octopus to Pacific:
```
2021-04-14 10:05:16.883145 I | ceph-object-controller: setting multisite settings for object store "my-store"
2021-04-14 10:05:17.006492 I | ceph-object-controller: Multisite for object-store: realm=my-store, zonegroup=my-store, zone=my-store
2021-04-14 10:05:17.006516 I | ceph-object-controller: multisite configuration for object-store my-store is complete
2021-04-14 10:05:17.006523 I | ceph-object-controller: creating object store "my-store" in namespace "rook-ceph"
2021-04-14 10:05:17.006533 I | cephclient: getting or creating ceph auth key "client.rgw.my.store.a"
2021-04-14 10:05:17.298176 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" started
2021-04-14 10:05:17.307913 I | ceph-object-controller: object store "my-store" deployment "rook-ceph-rgw-my-store-a" already exists. updating if needed
2021-04-14 10:05:17.317441 I | op-k8sutil: updating deployment "rook-ceph-rgw-my-store-a" after verifying it is safe to stop
2021-04-14 10:05:17.317463 I | op-mon: checking if we can stop the deployment rook-ceph-rgw-my-store-a
2021-04-14 10:05:25.552648 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-mgr-a"
2021-04-14 10:05:25.552662 I | op-mon: checking if we can continue the deployment rook-ceph-mgr-a
2021-04-14 10:05:25.555077 I | op-mgr: setting services to point to mgr "a"
2021-04-14 10:05:25.608535 I | op-osd: start running osds in namespace "rook-ceph"
2021-04-14 10:05:25.608587 I | op-osd: wait timeout for healthy OSDs during upgrade or restart is "10m0s"
2021-04-14 10:05:25.618392 I | op-osd: start provisioning the OSDs on PVCs, if needed
2021-04-14 10:05:25.620512 I | op-osd: no storageClassDeviceSets or volumeSources are defined to configure OSDs on PVCs
2021-04-14 10:05:25.620542 I | op-osd: start provisioning the OSDs on nodes, if needed
2021-04-14 10:05:25.628416 I | op-osd: 1 of the 1 storage nodes are valid
2021-04-14 10:05:25.760702 I | op-k8sutil: Removing previous job rook-ceph-osd-prepare-minikube to start a new one
2021-04-14 10:05:25.773213 I | op-k8sutil: batch job rook-ceph-osd-prepare-minikube still exists
2021-04-14 10:05:26.595673 I | op-mgr: successful modules: prometheus
2021-04-14 10:05:27.129569 I | op-mgr: successful modules: mgr module(s) from the spec
2021-04-14 10:05:27.776720 I | op-k8sutil: batch job rook-ceph-osd-prepare-minikube deleted
2021-04-14 10:05:27.781881 I | op-osd: started OSD provisioning job for node "minikube"
2021-04-14 10:05:27.786693 I | op-osd: OSD orchestration status for node minikube is "starting"
2021-04-14 10:05:28.617937 I | op-mgr: successful modules: balancer
2021-04-14 10:05:28.706307 W | cephclient: all OSDs are running on the same host. not performing upgrade check. running in best-effort
2021-04-14 10:05:28.710245 I | op-osd: updating OSD 0 on node "minikube"
2021-04-14 10:05:31.596511 I | op-mgr: the dashboard secret was already generated
2021-04-14 10:05:31.926125 I | op-mgr: setting ceph dashboard "admin" login creds
2021-04-14 10:05:32.190343 E | op-mgr: failed modules: "dashboard". failed to initialize dashboard: failed to set login credentials for the ceph dashboard: failed to set login creds on mgr: failed to complete command for set dashboard creds: Invalid command: unused arguments: ["P.c(b6V$0pfK#)'70c5z"]
dashboard set-login-credentials <username> : Set the login credentials. Password read from -i <file>
Traceback (most recent call last):
File "/usr/bin/ceph", line 1310, in <module>
retval = main()
File "/usr/bin/ceph", line 1256, in main
outf.write(outbuf)
TypeError: a bytes-like object is required, not 'str'
.
2021-04-14 10:05:35.665419 I | op-k8sutil: finished waiting for updated deployment "rook-ceph-rgw-my-store-a"
2021-04-14 10:05:35.665553 I | op-mon: checking if we can continue the deployment rook-ceph-rgw-my-store-a
2021-04-14 10:05:35.675293 I | ceph-object-controller: config map "rook-ceph-rgw-my-store-mime-types" for object store "my-store" already exists, not overwriting
2021-04-14 10:05:35.691549 I | ceph-object-controller: found more rgw deployments 3 than desired 3 in object store "my-store", scaling down
2021-04-14 10:05:35.691580 I | op-k8sutil: removing deployment rook-ceph-rgw-my-store-c if it exists
2021-04-14 10:05:35.698382 I | op-k8sutil: Removed deployment rook-ceph-rgw-my-store-c
2021-04-14 10:05:35.704568 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:05:44.881129 I | op-osd: OSD orchestration status for node minikube is "orchestrating"
2021-04-14 10:05:44.882149 I | op-osd: OSD orchestration status for node minikube is "completed"
2021-04-14 10:05:45.756762 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:05:45.782039 W | cephclient: all OSDs are running on the same host. not performing upgrade check. running in best-effort
2021-04-14 10:05:45.785188 I | op-osd: updating OSD 1 on node "minikube"
2021-04-14 10:05:54.768851 W | cephclient: all OSDs are running on the same host. not performing upgrade check. running in best-effort
2021-04-14 10:05:54.773191 I | op-osd: updating OSD 2 on node "minikube"
2021-04-14 10:05:55.843879 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:06:05.255136 I | op-osd: finished running OSDs in namespace "rook-ceph"
2021-04-14 10:06:05.255155 I | ceph-cluster-controller: done reconciling ceph cluster in namespace "rook-ceph"
2021-04-14 10:06:05.560201 W | ceph-cluster-controller: upgrade orchestration completed but somehow we still have more than one Ceph version running. map[ceph version 15.2.9 (357616cbf726abb779ca75a551e8d02568e15b17) octopus (stable):3 ceph version 16.2.0 (0c2054e95bcd9b30fdd908a79ac1d8bbc3394442) pacific (stable):6]:
2021-04-14 10:06:05.885001 I | op-k8sutil: "rook-ceph-rgw-my-store-c" still found. waiting...
2021-04-14 10:06:08.321209 I | exec: timeout waiting for process radosgw-admin to return. Sending interrupt signal to the process
2021-04-14 10:06:10.645672 I | ceph-spec: object "rook-ceph-rgw-my-store-c" matched on delete, reconciling
2021-04-14 10:06:11.940994 I | op-k8sutil: confirmed rook-ceph-rgw-my-store-c does not exist
2021-04-14 10:06:11.972935 I | ceph-spec: object "rook-ceph-rgw-my-store-c-keyring" matched on delete, reconciling
2021-04-14 10:06:11.974394 I | ceph-object-controller: deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-c"
2021-04-14 10:06:13.721964 I | ceph-object-controller: successfully deleted rgw config for "client.rgw.my.store.c" in mon configuration database
2021-04-14 10:06:13.721989 I | cephclient: deleting ceph auth "client.rgw.my.store.c"
2021-04-14 10:06:14.073700 I | ceph-object-controller: completed deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-c"
2021-04-14 10:06:14.073724 I | op-k8sutil: removing deployment rook-ceph-rgw-my-store-b if it exists
2021-04-14 10:06:14.083950 I | op-k8sutil: Removed deployment rook-ceph-rgw-my-store-b
2021-04-14 10:06:14.090226 I | op-k8sutil: "rook-ceph-rgw-my-store-b" still found. waiting...
2021-04-14 10:06:24.174846 I | op-k8sutil: "rook-ceph-rgw-my-store-b" still found. waiting...
2021-04-14 10:06:30.505574 I | ceph-spec: object "rook-ceph-rgw-my-store-b" matched on delete, reconciling
2021-04-14 10:06:32.201263 I | op-k8sutil: confirmed rook-ceph-rgw-my-store-b does not exist
2021-04-14 10:06:32.219559 I | ceph-object-controller: deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-b"
2021-04-14 10:06:32.220219 I | ceph-spec: object "rook-ceph-rgw-my-store-b-keyring" matched on delete, reconciling
2021-04-14 10:06:33.821969 I | ceph-object-controller: successfully deleted rgw config for "client.rgw.my.store.b" in mon configuration database
2021-04-14 10:06:33.821992 I | cephclient: deleting ceph auth "client.rgw.my.store.b"
2021-04-14 10:06:34.182091 I | ceph-object-controller: completed deleting rgw CephX key and configuration in centralized mon database for "rook-ceph-rgw-my-store-b"
2021-04-14 10:06:34.188659 I | ceph-object-controller: successfully scaled down rgw deployments to 1 in object store "my-store"
2021-04-14 10:06:34.188676 I | ceph-object-controller: enabling rgw dashboard
2021-04-14 10:06:49.487902 I | exec: timeout waiting for process radosgw-admin to return. Sending interrupt signal to the process
2021-04-14 10:06:49.494605 W | ceph-object-controller: failed to enable dashboard for rgw. failed to create user "dashboard-admin": failed to create s3 user: signal: interrupt
2021-04-14 10:06:49.494676 I | ceph-object-controller: created object store "my-store" in namespace "rook-ceph"
```
Running pods:
```
...
...
rook-ceph-rgw-my-store-a-6756bdbbd4-7265v 1/1 Running 1 41m
rook-ceph-rgw-my-store-a-6756bdbbd4-l5ccr 1/1 Running 1 41m
rook-ceph-rgw-my-store-a-6756bdbbd4-m679c 1/1 Running 1 41m
...
...
```
Signed-off-by: Sébastien Han <seb@redhat.com>
If the health spec of the CephCluster CRD has a timeout set to 0 like
so:
```
healthCheck:
daemonHealth:
mon:
disabled: false
interval: 45s
timeout: 0
```
And the mon goes out of quorum then Rook will not fail over the mon.
This is interesting when doing maintenance on a monitor and we don't
want to create a new one.
Signed-off-by: Sébastien Han <seb@redhat.com>
With Pacific comes the support for dualstack where ceph daemons can
listen on both ipv4 and ipv6 stacks.
A new field in the network spec has been added: `dualStack`
Signed-off-by: Sébastien Han <seb@redhat.com>
Update OSDs in parallel per the design in
design/ceph/update-osds-in-parallel.md
The max number of OSDs updated in parallel is currently fixed at 20.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
When a private docker registry is used and an image pull secret is specified in the chart, the pods with default Service Account fail to pull the image due to authentication issues.
Added rook-ceph-default service account and modify the pods specifications by adding the serviceAccountName.
Closes: https://github.com/rook/rook/issues/6673
Co-authored-by: Tareq Sharafy <tareq.sha@gmail.com>
Signed-off-by: parth-gr <partharora1010@gmail.com>
The mgr daemon may be failed over by ceph if the active mgr is not
responding and the standby mgr is available. If the active mgr changes
the services for the dashboard and metrics will be updated with a
label selector for the new active mgr. The services cannot direct
traffic to the standby mgr or else they will be incorrectly redirected
to the active mgr.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The GRPC metrics exposed by both the provisioner pod
and the node plugin pod on some port. Provisioner pod
is running on the pod network and the daemonset pods
run on the host network. sometimes starting the GRPC
metrics by default can lead to node plugin pods
crashloopback state this is due to the port conflict.
Moreover, the GRPC metrics are not for the user it's
for the one which will help to debug the time taken
by cephcsi to serve each GRPC call. Enabling it by
default won't be a good idea. So the plan is to
disable it by default, If someone faces any issue it
can be enabled later at some point in time.
closes#7378
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
For simple OSD scenario that don't require the following:
* an encrypted osd
* an osd with a backing device
* multiple osd on the same block device
The OSD will be prepared using the ceph-volume raw mode which does
not use LVM on non-pvc case.
Closes: https://github.com/rook/rook/issues/4768
Signed-off-by: Sébastien Han <seb@redhat.com>
The upcoming Rook release 1.6 will support Ceph Pacific which introduces
stable support for multiple Ceph Filesystems in the same cluster. So we
now allow the creation of multiple Filesystem if the release is Pacific.
Also, we remove the Operator config flag which was not necessary.
Here I have 2 filesystems:
```
[root@rook-ceph-tools-6b4889fdfd-vncck /]# ceph fs ls
name: myfs, metadata pool: myfs-metadata, data pools: [myfs-data0 ]
name: myfs2, metadata pool: myfs2-metadata, data pools: [myfs2-data0 ]
```
```
cluster:
id: cc747950-589a-4862-bad2-ff7de1217b1a
health: HEALTH_WARN
mons a,b,c are low on available space
6 pool(s) have no replicas configured
services:
mon: 3 daemons, quorum a,b,c (age 2d)
mgr: a(active, since 2d)
mds: myfs:1 myfs2:1 {myfs2:0=myfs2-b=up:active,myfs:0=myfs-b=up:active} 2 up:standby-replay
osd: 3 osds: 3 up (since 2m), 3 in (since 6d)
data:
pools: 6 pools, 717 pgs
objects: 44 objects, 4.2 KiB
usage: 46 MiB used, 90 GiB / 90 GiB avail
pgs: 717 active+clean
io:
client: 1.8 KiB/s rd, 4 op/s rd, 0 op/s wr
progress:
Global Recovery Event (45s)
[===========================.]
```
Signed-off-by: Sébastien Han <seb@redhat.com>
Remove features and design that supports adding OSDs to Ceph clusters
via `spec:driveGroups`. Update the Ceph upgrade doc that informs users
who currently use Drive Groups (we believe there are none of these
users) how to migrate to using the `spec:storage` config.
Resolves https://github.com/rook/rook/issues/7275
Revert "ceph: fix drive group deployment failure"
This reverts commit 76f1d9944e.
Revert "ceph: osd: add drive groups spec to cluster CR"
This reverts commit 7117fc12b7.
Revert "design: ceph orchestrator module add/remove OSDs"
This reverts commit 178187d035.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.
The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050
The same goes for integration tests, they will get added later once we
start testing on Pacific.
Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>
The PDBs have been stabilizing for a good length of time now since
they were created in v1.1, and later redesigned in v1.5. The time
has come to enable the feature by default so everyone can enjoy
the stable Rook storage even while draining K8s nodes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The KMS encryption via HashiCorp Vault can be consumed for RGW, adding those details
in the doc.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
This was the last remaining CRD to not use the controller-runtime
library.
Small additions were added with the transition:
* the Kubernetes Secret that contains the CephX key has now an owner
reference to the CephClient object
* the secret name is present in the Status field of the CephClient:
```
status:
info:
secretName: rook-ceph-client-glance
phase: Ready
```
The controller will reconcile on CR updates and also if the Kubernetes
Secret is deleted.
Closes: https://github.com/rook/rook/issues/4938
Signed-off-by: Sébastien Han <seb@redhat.com>
Now, we can schedule snapshots on pools from the CephBlockPool CR when
the pool is mirrored.
It can be enabled like this:
```
mirroring:
enabled: true
mode: pool
snapshotSchedules:
- interval: 24h # daily snapshots
startTime: 14:00:00-05:00
```
Multiple schedules are supported since snapshotSchedules is a list.
Signed-off-by: Sébastien Han <seb@redhat.com>
This updates the chart to make use of helm3 which has been released
for some time, and also permits CRDs to be installed pror to other objects
allowing the chart to be deployed at the same time as CRs for rook objects.
Co-authored-by: Pete Birley <pete@port.direct>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The discovery daemon is not needed in most scenarios, therefore we disable it
by default. More and more clusters are moving to the cluster-on-pvc scenario
which certainly does not need the local discovery. Even where clusters are not
running on PVCs, the discovery is not needed since the device discovery is again
performed in the osd prepare job.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>