The KMS encryption via HashiCorp Vault can be consumed for RGW, adding those details
in the doc.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
This was the last remaining CRD to not use the controller-runtime
library.
Small additions were added with the transition:
* the Kubernetes Secret that contains the CephX key has now an owner
reference to the CephClient object
* the secret name is present in the Status field of the CephClient:
```
status:
info:
secretName: rook-ceph-client-glance
phase: Ready
```
The controller will reconcile on CR updates and also if the Kubernetes
Secret is deleted.
Closes: https://github.com/rook/rook/issues/4938
Signed-off-by: Sébastien Han <seb@redhat.com>
Now, we can schedule snapshots on pools from the CephBlockPool CR when
the pool is mirrored.
It can be enabled like this:
```
mirroring:
enabled: true
mode: pool
snapshotSchedules:
- interval: 24h # daily snapshots
startTime: 14:00:00-05:00
```
Multiple schedules are supported since snapshotSchedules is a list.
Signed-off-by: Sébastien Han <seb@redhat.com>
This updates the chart to make use of helm3 which has been released
for some time, and also permits CRDs to be installed pror to other objects
allowing the chart to be deployed at the same time as CRs for rook objects.
Co-authored-by: Pete Birley <pete@port.direct>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The discovery daemon is not needed in most scenarios, therefore we disable it
by default. More and more clusters are moving to the cluster-on-pvc scenario
which certainly does not need the local discovery. Even where clusters are not
running on PVCs, the discovery is not needed since the device discovery is again
performed in the osd prepare job.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Due to #6492, preservePoolsOnDelete is not useful at all for CephFS;
after the filesystem is deleted, the leftover pools cannot be
reassocaited with a newly created filesystem without wiping all
metadata. The only way we can actually preserve data is keeping around
the entire filesystem.
This commit implements a `preserveFilesystemOnDelete` option which work
similar to the existing pool preservation option but instead keeps the
whole CephFS while taking it down and removing all MDSes.
This commit also changes all documentation to refer to this new option
with the intent of essentially deprecating `preservePoolsOnDelete`. IMO,
keeping around `preservePoolsOnDelete` is actively harmful because it
lulls users into thinking their data will be safe but, in reality,
recovering from this situation is highly complex and has large potential
for data loss.
Signed-off-by: Lalit Maganti <lalitm@google.com>
When the Ceph cluster runs on PVC and the OSDs are encrypted we can
store LUKS's Key Encryption Key inside a Key Management System. Today,
Rook only supports HashiCorp Vault: https://www.vaultproject.io/
The CephCluster has now a new "security" field which will plug onto the
KMS. Here is an example:
security:
kms:
tokenSecretName: <name of the secret containing a Vault token, used
to authenticate>
connectionDetails: < a map of strings containing connection
information>
Refer to the ceph-cluster-crd documentation to lear more.
Closes: https://github.com/rook/rook/issues/6105
Signed-off-by: Sébastien Han <seb@redhat.com>
the apiextensions.k8s.io/v1beta1 version of CustomResourceDefinition
is deprecated in Kubernetes v1.16 and will no longer be supported from
v1.19. For now, we changing only for ceph and it's related documented.
Signed-off-by: subhamkrai <srai@redhat.com>
Co-authored-by: travisn <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
this commit update PendingReleaseNotes.md with
PR 6475 i.e export the storage capacity of
the ceph cluster.
Signed-off-by: subhamkrai <srai@redhat.com>
The pool spec has now a new property called "replicasPerFailureDomain"
which essentially represents the number of replicas to store in each
failure domain.
Assuming the failure domain is a datacenter (if the cluster is
stretched) then you will have 2 replicas per datacenter where each
replica ends up on a different host. This gives you a total of 4
replicas and for this, the "size" must be set to 4.
Closes: https://github.com/rook/rook/issues/5591
Signed-off-by: Sébastien Han <seb@redhat.com>
The ceph mons should never be started with an even number in quorum.
The desired state should always be an odd number of mons to ensure
a healthy majority quorum. The operator now rejects a request for
an even number of mons.
If there are not a sufficient number of nodes for mons to be on a
unique host, the operator also returns an early error instead of
getting stuck waiting for a pending mon. This is not foolproof since
there might be fewer available nodes for the mons, but at least it
is a quick check for the common case.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook is now capable of configuring mirroring between sites. The
implementation works at different levels:
* CephBlockPool: which introduces a new `mirroring` configuration as well
as `statusCheck`. When turned on, Rook will enable mirroring on the
pool. It will also create a bootstrap peer token and store it in a
Kubernetes Secret. The name of that Secret can be found in the Status
field of the CephBlockPool CRD. This token can be fetched and used by
other clusters to configure the site as a peer. Mirroring can be
configured either at the pool or the image level.
* CephRBDMirror: which introduces a new `peers` configuration allowing
Rook to connect to peers by passing a Secret name. The administrator will
create a Kubernetes Secret with 2 keys: 'token' for the bootstrap peer
token and 'pool' for the name of pool. Once detected the rbd-mirror
controller will go ahead and import the peer configuration.
Pool mirroring status example:
```
status:
info:
rbdMirrorBootstrapPeerSecretName: pool-peer-token-test
mirroringInfo:
lastChanged: "2020-09-17T14:47:27Z"
lastChecked: "2020-09-17T14:48:27Z"
summary:
summary:
mode: image
peers:
- client_name: client.rbd-mirror-peer
direction: rx-tx
mirror_uuid: ""
site_name: rhcs
uuid: c50522a4-28a4-4bd3-ba68-e11780308882
site_name: 91eae0dd-06b1-4d2c-91f3-1311c9df382b-rook-ceph
mirroringStatus:
lastChecked: "2020-09-17T14:48:27Z"
summary:
summary:
daemon_health: OK
health: OK
image_health: OK
states:
replaying: 1
```
Signed-off-by: Sébastien Han <seb@redhat.com>
We can now encrypted OSD device that were provisioned via a storage
class using the PV interface.
The encryption works at the storageClassDeviceSets level, which means we
can have encrypted and non-encrypted sets.
Using the new key `encrypted` we can turn it on.
Signed-off-by: Sébastien Han <seb@redhat.com>
A long-standing issue that was only solved via a documentation note. Now,
during the prepare pod instantiation, we check for the presence of the
lvm binary on the host. Success will indicate that we can bootstrap
OSDs where failure will refuse to prepare the OSD.
Closes: https://github.com/rook/rook/issues/5627
Signed-off-by: Sébastien Han <seb@redhat.com>
The Cassandra sidecar uses the Jolokia javaagent to expose a REST API of
Cassandra's administrative interface, which is normally only exposed via
JMX, a Java-only RPC protocol. Update the Jolokia javaagent to a newer
version, containing new features and security bug fixes.
Signed-off-by: Yannis Zarkadas <yanniszark@arrikto.com>
Pass desired devices to the OSD provisioning container by
JSON-marshalling/-unmarshalling a disk ID with StorageConfig settings.
This will allow disks to be specified that contain special characters.
Notably, this will support /dev/disk/by-path/pci-HHHH:HH:HH.H devices
that were unsupported previously due to the format which separated
devices from params with colons (:).
Fixes#5056Fixes#5535
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
We can now connect an external prometheus exporter to Rook to collect
metrics and generate alerts from Prometheus.
Just enable this in the CephCluster CR:
```
spec:
external:
enable: true
monitoring:
enabled: true
rulesNamespace: rook-ceph
externalMgrEndpoints:
- ip: 192.168.39.182
```
Closes: https://github.com/rook/rook/issues/5516
Signed-off-by: Sébastien Han <seb@redhat.com>
Add the ability to provision Ceph OSDs with Drive Groups.
This adds Drive Groups to the CephCluster CRD, and it sets code
in place for propagating this config to the OSD provisioning pod.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
This commit allows us to configure status check for each daemon:
* "mon": health check on the ceph monitors (quorum)
* "osd": health check on the ceph osds
* "status": ceph health status check
Each check is controlled by the following settings:
* disabled: whether to disable the check (default: false)
* internal: interval to run the check
* timeout: only valid for mons, is the timeout for unresponsive mon
before failling over.
Example to disable the status health check:
```yaml
healthCheck:
daemonHealth:
status:
disabled: true
```
As part of that, pod's livenessprobe can now be configured via the
following settings:
* disabled: whether to enable or not
* probe: override the current probe in place by a new one
```yaml
healthCheck:
livenessProbe:
mon:
disabled: true
```
Closes: https://github.com/rook/rook/issues/5772
Signed-off-by: Sébastien Han <seb@redhat.com>
Now, when an S3 user gets created, Rook will add the S3 endpoint to the
Secret along with the credentials.
Signed-off-by: Sébastien Han <seb@redhat.com>
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.
A good status will look like:
status:
endpointStatus:
lastChanged: "2020-06-25T13:47:45Z"
lastChecked: "2020-06-25T13:48:46Z"
phase: Connected
A failed status:
status:
endpointStatus:
details: |-
error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
health: ERROR
This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.
Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
adds deploy.sh script to deploy validatingwebhookconfiguration and create secrets.
adds new command ceph admission-controller to start webhook servers.
adds validation for various rook custom resources
Signed-off-by: Vineet Badrinath <vbadrina@redhat.com>
We can now explicitly set any property on a given pool by using the new
Property field in the CephBlockPool Spec.
Also, this fixes the case where both `CephBlockPool` and `CephCluster`
are created at the same time. When Rook creates the pool, the cluster is
still being bootstrapped and the global option
`osd_pool_default_pg_autoscale_mode` has not bee set yet. So the pool
gets created but its `pg_autoscale_mode` property is set to `warn`
instead of `on`.
Closes: https://github.com/rook/rook/issues/5608V
Signed-off-by: Sébastien Han <seb@redhat.com>
OSD on PVC supports multipath thanks to the following commit.
ceph: add support for multipath devices
32730123ea
However, it's not documented yet.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.
Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
As another option to run Ceph commands, a script can be executed
with Ceph commands as if running in the toolbox, but will instead
run in a job where the logs can be collected and analyzed
separately. This would be useful where automation is in place to collect
information about the cluster periodically or upon some failure
that is detected, without requiring an interactive shell.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Previously, the rbd-mirror daemon was integrated into the `CephCluster`
CRD. This wasn't really practical since we would have to wait for the
whole orchestration to be done to actually set it up. The same goes for
any CR update. Let's say you want to change the number of daemons, Rook
would go through mons, mgrs and osds until it get to rbd-mirror.
This triggers an undesired full orchestration.
With its own CRD this component just gains a lot more flexibility.
Closes: https://github.com/rook/rook/issues/5084
Signed-off-by: Sébastien Han <seb@redhat.com>
You can now use Rook along with Multus. Multus must be up and running
and the right ressources must exist such as NetworkAttachmentDefinition
CR.
The Cluster CR spec has new fields to work with multus:
network:
provider: multus (or 'host' for hostNetworking)
selectors:
public: NetworkAttachmentDefinition name
cluster: NetworkAttachmentDefinition name
If only a single NetworkAttachmentDefinition is provided Rook will use
both anyway for the Ceph traffic.
Please refer to the doc to learn more.
Closes: https://github.com/rook/rook/issues/4716
Signed-off-by: Sébastien Han <seb@redhat.com>
Various updates are needed for the v1.3 release.
- The pending release notes were missing some features
- Added a section to the upgrade guide for breaking changes to OSDs
- Clarifications around Ceph prereqs and min version
- Other misc clarifications
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Thanks to https://github.com/ceph/ceph/pull/33633, when running on
Octopus, we don't need to use `--pid=host`. This means we are not using
the host PID but use the PID namespace of the pod.
Signed-off-by: Sébastien Han <seb@redhat.com>