If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.
Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
The first patch to configure vault for ceph object store. If the `security.kms` configured in
`clusterSpec` CRD, RGW will be configured with vault kms settings to handle SSE request from s3 clients.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
By setting the working directory of a Ceph daemon to the log direcor
(which is bindmounted to the host), we can ensure that the coredumps
will be available.
On CentOS 7 **only**, the kernel is configured with:
```
cat /proc/sys/kernel/core_pattern
core
```
which means that the coredumps will end up in the process working
directory and will be named "core".
On CentOS 8, everything is different and this won't work but still
CentOS will be able to consume this and changing the working directory
is harmless.
Signed-off-by: Sébastien Han <seb@redhat.com>
If the object store is not found during deletion of the object store CR,
proceed with the deletion instead of blocking and re-queueing the
deletion reconcile in an endless loop. This was causing instability
in the integration tests.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
DeleteUser was returning error "failed to delete user ... with buckets"
without corresponding buckets information. Now it will check buckets, If
buckets are present then return error with buckets info on failure.
Signed-off-by: Nitin Goyal <nigoyal@redhat.com>
After #6217, internal rgw spawning on external cluster was still broken,
operator claiming:
```
ceph-object-controller: failed to reconcile failed to create object
store deployments: failed to reconcile external endpoint:
failed to create or update object store "arch-cloud" endpoint:
failed to create endpoint "rook-ceph-rgw-arch-cloud". Endpoints
"rook-ceph-rgw-arch-cloud" is invalid: subsets[0]:
Required value: must specify `addresses` or `notReadyAddresses`
```
To spawn internal rgw for external cluster, changes detection of
what mean 'internal' for rgw and only use the externalRgwEndpoints
list for that independently of the status of the cluster
Now there is multiple posibilities:
* internal cluster and internal rgw pods: "normal case"
* external cluster and internal rgw pods: <= now working with the PR
* external cluster and external rgw pods: the external case
* internal cluster and external rgw pods: <= new case that could exist
Signed-off-by: Julien Girardin <jugirardin@free.fr>
The command to set or disable the rgw dashboard started
hanging in some scenarios in v15.2.8. For now we start
the rgw dashboard config in a goroutine until this issue
is tracked down. After the issue is fixed in ceph, the
goroutines will no longer be necessary.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Update to the latest lib bucket provisioner code.
Fixes issue 6650
Modifies CRD for objectbucketclaims to fix an additional bug where an
ObjectBucket's 'ClaimRef' is lost due to the CRD validation being
specified incorrectly.
Changes OBC deletion/cleanup to delete the bucket before the user. A
user cannot be deleted without an unsafe purge option if the user has
buckets associated to it.
Does not reintroduce bug 6767 from previous fix for 6650
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The legacy rgw deployments were removed in 1.0, so there is no more need
for the operator to keep checking for those!
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
If the selector changes on the rgw service, during upgrade the
clients will briefly not be able to connect to the rgw pods.
After all the pods are updated with any new labels, the connections
would be restored, but we want to avoid that temporary outage.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We can now collect logs directly into a side-car container.
A new CRD spec has been added:
spec:
logCollector:
enabled: true
periodicity: 24h
Every 24h we will rotate log files for each Ceph daemon.
Signed-off-by: Sébastien Han <seb@redhat.com>
We must add the finalizer right after the object creation otherwise the
seerver will later return an error on update that the object has been
modified. Indeed, it has been by the task that updates the status when
the object is first created.
Signed-off-by: Sébastien Han <seb@redhat.com>
The RGW deployment's ceph-version label should now display the same
version as the image it is using.
It previously displayed an empty version: 0.0.0-0.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
Pools in stretch clusters will all use the same crush rule that
is generated by the operator when configuring the cluster for
stretch mode, with replica 4 and two replicas per failure domain.
Pools cannot create new crush rules, so we require that replica be
4 when the pools is created, EC is not allowed, and no new
crush rule will be created.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.
By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.
To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.
The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.
Fixes#4993.
Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>
Add unit tests for when the "Zone" is configured
in an object store so that the object store joins
the CephObjectZone and the zone's corresponding
multisite configuration.
Signed-off-by: Ali Maredia <amaredia@redhat.com>
The pool spec has now a new property called "replicasPerFailureDomain"
which essentially represents the number of replicas to store in each
failure domain.
Assuming the failure domain is a datacenter (if the cluster is
stretched) then you will have 2 replicas per datacenter where each
replica ends up on a different host. This gives you a total of 4
replicas and for this, the "size" must be set to 4.
Closes: https://github.com/rook/rook/issues/5591
Signed-off-by: Sébastien Han <seb@redhat.com>
When rgw is not responding the aws-sdk retries using exponential
backoff. Having 20 retries means we would wait for a while (hours) to
eventually fail. This is a problem as we cannot reflect the status of
the bucket properly.
Now we use 5 max retries which means around 8 seconds.
Signed-off-by: Sébastien Han <seb@redhat.com>
During the 1.3 cycle, we had implemented support for OBC external mode
through an Endpoint in the StorageClass parameter. In 1.4, this is not
the case anymore as we directly use the CephObjectStore external mode.
However, we must maintain backward compatibility with the cluster using
the Endpoint and not fail the upgrade.
Thus checking for a CephObjectStore is invalid and should not always be
assumed, now if we detect an Endpoint we return a simple context which
fixes the upgrade.
Signed-off-by: Sébastien Han <seb@redhat.com>
this commit handle golangci-lint linter errcheck.
`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases
To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`
Signed-off-by: subhamkrai <srai@redhat.com>
User unlink is performed just before deleting OBC resources.
Since user is internal to OBC and next to remove to delete user,
unlinking user is not necessary.
Fixes: 6300
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
this commit handle golangci-lint linter staticcheck error.
`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.
To see only `staticcheck` linter output
`golangci-lint run --disable-all -E staticcheck`
Signed-off-by: subhamkrai <srai@redhat.com>
this commit will enable one more linter ineffassign
in golangci-lint.
This linter throws an error when variable is assigned and never used.
`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>