If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.
Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
The first patch to configure vault for ceph object store. If the `security.kms` configured in
`clusterSpec` CRD, RGW will be configured with vault kms settings to handle SSE request from s3 clients.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
In some cases the user dont want to run snapshotter
container either for CephFS or RBD. In that case the
user wont install the required snapshot CRD's due
to that the snapshotter sidecar container produces
lot of noisy logs.
Snapshotter will be enabled by default for both
CephFS and RBD, but with this PR we are providing
an option to disable snapshotter sidecar deployment
either for CephFS or RBD.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Multiple OSD pods for the same OSD might run simultaneously becasue OSD pod
is managed by Deployment resource. However, OSD locking mechanism
doens't work because the lock file (fsid file) exist for each OSD pod.
It resutls in OSD corruption.
Closes: https://github.com/rook/rook/issues/6530
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
This was the last remaining CRD to not use the controller-runtime
library.
Small additions were added with the transition:
* the Kubernetes Secret that contains the CephX key has now an owner
reference to the CephClient object
* the secret name is present in the Status field of the CephClient:
```
status:
info:
secretName: rook-ceph-client-glance
phase: Ready
```
The controller will reconcile on CR updates and also if the Kubernetes
Secret is deleted.
Closes: https://github.com/rook/rook/issues/4938
Signed-off-by: Sébastien Han <seb@redhat.com>
By setting the working directory of a Ceph daemon to the log direcor
(which is bindmounted to the host), we can ensure that the coredumps
will be available.
On CentOS 7 **only**, the kernel is configured with:
```
cat /proc/sys/kernel/core_pattern
core
```
which means that the coredumps will end up in the process working
directory and will be named "core".
On CentOS 8, everything is different and this won't work but still
CentOS will be able to consume this and changing the working directory
is harmless.
Signed-off-by: Sébastien Han <seb@redhat.com>
If the object store is not found during deletion of the object store CR,
proceed with the deletion instead of blocking and re-queueing the
deletion reconcile in an endless loop. This was causing instability
in the integration tests.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If ssl was not enabled, the dashboard was being restarted with every
reconcile. Instead, the dashboard should only be restarted when
the settings have changed in order to reduce the frequency that
the dashboard is restarted.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
DeleteUser was returning error "failed to delete user ... with buckets"
without corresponding buckets information. Now it will check buckets, If
buckets are present then return error with buckets info on failure.
Signed-off-by: Nitin Goyal <nigoyal@redhat.com>
When the version is detected it will print the name of the release
Pacific correctly where previously it was showing <unknown version>".
Signed-off-by: Sébastien Han <seb@redhat.com>
`.spec.resources.osd` resources and `placement: all:` placements
parameters for the OSD pod did not reflect the value like
it has for other pods mon,mgr.
this commit will allow to set those parameters for the OSD pod also.
Signed-off-by: subhamkrai <srai@redhat.com>
After #6217, internal rgw spawning on external cluster was still broken,
operator claiming:
```
ceph-object-controller: failed to reconcile failed to create object
store deployments: failed to reconcile external endpoint:
failed to create or update object store "arch-cloud" endpoint:
failed to create endpoint "rook-ceph-rgw-arch-cloud". Endpoints
"rook-ceph-rgw-arch-cloud" is invalid: subsets[0]:
Required value: must specify `addresses` or `notReadyAddresses`
```
To spawn internal rgw for external cluster, changes detection of
what mean 'internal' for rgw and only use the externalRgwEndpoints
list for that independently of the status of the cluster
Now there is multiple posibilities:
* internal cluster and internal rgw pods: "normal case"
* external cluster and internal rgw pods: <= now working with the PR
* external cluster and external rgw pods: the external case
* internal cluster and external rgw pods: <= new case that could exist
Signed-off-by: Julien Girardin <jugirardin@free.fr>
After the operator restarts, the controllers will not all be able
to reconcile until the ceph config has been generated by the reconcile
of the CephCluster controller. The message printed to the operator log
is frequently seen as an error condition even though it is a normal
condition where we requeue the reconcile until the config is available.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The command to set or disable the rgw dashboard started
hanging in some scenarios in v15.2.8. For now we start
the rgw dashboard config in a goroutine until this issue
is tracked down. After the issue is fixed in ceph, the
goroutines will no longer be necessary.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Update to the latest lib bucket provisioner code.
Fixes issue 6650
Modifies CRD for objectbucketclaims to fix an additional bug where an
ObjectBucket's 'ClaimRef' is lost due to the CRD validation being
specified incorrectly.
Changes OBC deletion/cleanup to delete the bucket before the user. A
user cannot be deleted without an unsafe purge option if the user has
buckets associated to it.
Does not reintroduce bug 6767 from previous fix for 6650
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The legacy rgw deployments were removed in 1.0, so there is no more need
for the operator to keep checking for those!
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator namespace is not used in the StartOperatorSettingsWatch()
method. Instead, the method looks up the namespace from the
POD_NAMESPACE env var.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Currently, mgr deployment has two overlapping environment variables, ROOK_POD_IP.
This causes kube-apiserver to create error logs
Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
Currently the reconcile of the clusterDisruption controller was triggered mainly due to the
ceph status update in the cephCluster CR. This PR adds reconciles the cluster when:
1. Reconcile when the cluster is created. (This will trigger the first reconcile)
2. Reconcile only when the clusterSpec is updated. (This will avoid triggers when cluster status is updated)
3. Reconcile for events on cephblockpool, cephfilesystem and cephObjectStore.
4. Reconcile for events on Main PDB and when `DisruptionsAllowed` is 0. (that is, when one of the OSD goes down).
5. Reconcile after 30 seconds when there is an active drain going on, that is, pdbStateMap has `draining-failure-domain` as not empty.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
The OSD configuration on local devices was being skipped if a
storageClassDeviceSet was specified. Both types of OSDs should
be allowed in the same cluster, which allows a cluster in the cloud
to also take advantage of local devices with different perf
properties.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The deviceClass property was being ignored when creating the
non-pvc OSDs. Now the deviceClass will be specified as a property
for individual devices, all devices on a node, or all OSDs in the
cluster, depending on the level where the config is applied in the
cluster CR.
Co-authored-by: shenjiatong <yshxxsjt715@gmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The arbiter can only be configured with the stretch cluster if the
CRUSH map is balanced and there are two zones in the CRUSH map.
After the OSDs are configured, we wait for all the OSD pods to be
running and that the CRUSH map is balanced. If it takes more than
two minutes, we fail the reconcile and try again. This is only done
the first time the stretch cluster is configured. In future reconciles
we first check if the stretch cluster is already enabled before
enabling it again.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The mon timeouts and pod retries were removed back in v1.1,
this just removes some obsolete variables from the legacy
mon scheduling.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator schedules all the mons at the same time, then
serially starts each of the mons. If the operator was restarted
or otherwise failed before all the mons were started, the next
reconcile will schedule new mons again and ignore the scheduling
that was already computed. Since the last mon running is always
expected to be in quorum, we can reset the maxMonId to be
the highest mon currently created.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If the kernel doesnt support mapping of
rbd image with deep-flatten feature,cephcsi
need to flatten the rbd image first and than
map the image on the node.
cephcsi first tries to add a task to flatten
the rbd image if its receives any permission
error it will try to call rbd CLI command which
is a blocking call.
This commit adds mgr caps to the csi-rbd-node user
so that cephcsi will add flatten task and return
immediate error to the kubelet, let kubelet retry
again,If we go with blocking rbd CLI call we may
end up having stale maps on the node in corner cases.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>