Commit Graph
286 Commits
Author SHA1 Message Date
Sébastien Han 28c78b9204 Merge pull request #7044 from leseb/fix-7022
ceph: update rgw and mds deployment for logCollector
2021-01-25 09:29:40 +01:00
Sébastien Han ebbf332d8d ceph: update rgw and mds deployment for logCollector
If the CephCluster CR spec is updated to activate the logCollector, we
must reflect that change onto child CRDs, like the mds and rgw since
their configurationn would be impacted too.
Now we watch for the CephCluster object changes from the object/file
controllers and react upon the appropriate event.

Closes: https://github.com/rook/rook/issues/7022
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-22 18:19:07 +01:00
Jiffin Tony Thottan da61c9a83e ceph: vault kms configuration for ceph object store
The first patch to configure vault for ceph object store. If the `security.kms` configured in
`clusterSpec` CRD, RGW will be configured with vault kms settings to handle SSE request from s3 clients.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-01-22 10:44:03 +05:30
Sébastien Han 093ed7dd62 Merge pull request #6984 from leseb/core-dumps-dir
ceph: set process working dir to /var/log/ceph
2021-01-20 10:39:42 +01:00
Sébastien Han cb7d600841 ceph: set process working dir to /var/log/ceph
By setting the working directory of a Ceph daemon to the log direcor
(which is bindmounted to the host), we can ensure that the coredumps
will be available.

On CentOS 7 **only**, the kernel is configured with:

```
cat /proc/sys/kernel/core_pattern
core
```

which means that the coredumps will end up in the process working
directory and will be named "core".

On CentOS 8, everything is different and this won't work but still
CentOS will be able to consume this and changing the working directory
is harmless.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-15 16:06:42 +01:00
Travis Nielsen ea29665a8a ceph: during object store deletion return success if not found
If the object store is not found during deletion of the object store CR,
proceed with the deletion instead of blocking and re-queueing the
deletion reconcile in an endless loop. This was causing instability
in the integration tests.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-13 15:44:28 -07:00
Sébastien Han 22721f1df8 Merge pull request #6949 from leseb/bump-controller-runtime-0.7
core: bump to controller-runtime 0.7.0 version
2021-01-13 17:16:13 +01:00
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
Nitin Goyal 76a370a57a ceph: enhance delete cephObjectStoreUser logging
DeleteUser was returning error "failed to delete user ... with buckets"
without corresponding buckets information. Now it will check buckets, If
buckets are present then return error with buckets info on failure.

Signed-off-by: Nitin Goyal <nigoyal@redhat.com>
2021-01-12 23:30:57 +05:30
Julien Girardin d63d9b7d4d ceph: change external rgw detection, not relying on cluster
After #6217, internal rgw spawning on external cluster was still broken,
operator claiming:
```
ceph-object-controller: failed to reconcile failed to create object
store deployments: failed to reconcile external endpoint:
failed to create or update object store "arch-cloud" endpoint:
failed to create endpoint "rook-ceph-rgw-arch-cloud". Endpoints
"rook-ceph-rgw-arch-cloud" is invalid: subsets[0]:
Required value: must specify `addresses` or `notReadyAddresses`
```

To spawn internal rgw for external cluster, changes detection of
what mean 'internal' for rgw and only use the externalRgwEndpoints
list for that independently of the status of the cluster
Now there is multiple posibilities:

 * internal cluster and internal rgw pods: "normal case"
 * external cluster and internal rgw pods: <= now working with the PR
 * external cluster and external rgw pods: the external case
 * internal cluster and external rgw pods: <= new case that could exist

Signed-off-by: Julien Girardin <jugirardin@free.fr>
2021-01-08 11:28:41 +01:00
Travis Nielsen 464e332dc1 ceph: init rgw dashboard access key in goroutine
The command to set or disable the rgw dashboard started
hanging in some scenarios in v15.2.8. For now we start
the rgw dashboard config in a goroutine until this issue
is tracked down. After the issue is fixed in ceph, the
goroutines will no longer be necessary.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-07 07:08:20 -07:00
Travis Nielsen 423bc3167d Merge pull request #6856 from travisn/obsolete-legacy-rgw-cleanup
ceph: No need for removal of legacy rgw deployments
2020-12-18 14:32:17 -07:00
Blaine Gardner b8dc2a3214 ceph: update lib bucket provisioner
Update to the latest lib bucket provisioner code.
Fixes issue 6650

Modifies CRD for objectbucketclaims to fix an additional bug where an
ObjectBucket's 'ClaimRef' is lost due to the CRD validation being
specified incorrectly.

Changes OBC deletion/cleanup to delete the bucket before the user. A
user cannot be deleted without an unsafe purge option if the user has
buckets associated to it.

Does not reintroduce bug 6767 from previous fix for 6650

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-18 12:48:54 -07:00
Travis Nielsen e8ea6dc2f8 ceph: no need for removal of legacy rgw deployments
The legacy rgw deployments were removed in 1.0, so there is no more need
for the operator to keep checking for those!

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-18 12:03:50 -07:00
Blaine Gardner 3fe61c115e ceph: add a missed cleanup item in obc provisioner
Add a cleanup item to remove user if it fails on S3 agent creation.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-04 10:47:32 -07:00
Blaine Gardner 22cd1bf0f2 Revert "ceph: update object bucket provisioner library"
This reverts commit 3b4ee6c8e3.

Revert the fix that introduces the following bug on upgrade:
https://github.com/rook/rook/issues/6767

This fix as well as a fix for the upgrade case will follow in days to
come.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-04 10:04:34 -07:00
Blaine Gardner bc08f51f65 Merge pull request #6699 from BlaineEXE/update-lib-bucket-provisioner
ceph: update object bucket provisioner library
2020-12-03 14:47:44 -07:00
Blaine Gardner 3b4ee6c8e3 ceph: update object bucket provisioner library
The library for object bucket provisioning is updated to fix errors
during provisioning.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-03 09:43:12 -07:00
Sébastien Han f27b954963 Merge pull request #6742 from travisn/rgw-service-selector
ceph: RGW service selector should not change during upgrade
2020-12-03 09:30:18 +01:00
Travis Nielsen 8913a14ce4 ceph: rgw service selector should not change
If the selector changes on the rgw service, during upgrade the
clients will briefly not be able to connect to the rgw pods.
After all the pods are updated with any new labels, the connections
would be restored, but we want to avoid that temporary outage.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-02 16:45:02 -07:00
Blaine Gardner e75dcac8a3 ceph: add more debug logging to obj bucket claims
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-01 14:01:33 -07:00
Sébastien Han c6a87203ca ceph: add log collector
We can now collect logs directly into a side-car container.
A new CRD spec has been added:

spec:
  logCollector:
    enabled: true
    periodicity: 24h

Every 24h we will rotate log files for each Ceph daemon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-12-01 16:30:17 +01:00
Sébastien Han 97be23e374 ceph: apply finalizer before updating object status
We must add the finalizer right after the object creation otherwise the
seerver will later return an error on update that the object has been
modified. Indeed, it has been by the task that updates the status when
the object is first created.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-18 18:16:59 +01:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Blaine Gardner 98d5c73d5e ceph: fill in rgw deployment version
The RGW deployment's ceph-version label should now display the same
version as the image it is using.

It previously displayed an empty version: 0.0.0-0.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-11-11 15:13:31 -07:00
Sébastien Han e3d032dbc1 Merge pull request #6489 from travisn/stretch-mons
ceph: Stretch cluster configuration
2020-11-06 15:16:39 +01:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
Travis Nielsen e72b69f8e0 ceph: validate pools for stretch clusters
Pools in stretch clusters will all use the same crush rule that
is generated by the operator when configuring the cluster for
stretch mode, with replica 4 and two replicas per failure domain.
Pools cannot create new crush rules, so we require that replica be
4 when the pools is created, EC is not allowed, and no new
crush rule will be created.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-05 09:23:57 -07:00
Travis Nielsen b91f4211c9 ceph: configure a stretched cluster
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-05 09:17:52 -07:00
Jonas Schäfer 7ace6ac255 ceph: make root= CRUSH label value configurable
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.

By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.

To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.

The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.

Fixes #4993.

Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>
2020-10-29 09:22:41 +01:00
subhamkrai cfdc5e7c50 ceph: remove allNode option for rgw
this commit removes unnecessary allNodes option
from rgw.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-10-28 22:38:11 +05:30
Ali Maredia 8964941d44 ceph: add unit tests for object store controller with multisite
Add unit tests for when the "Zone" is configured
in an object store so that the object store joins
the CephObjectZone and the zone's corresponding
multisite configuration.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-10-22 11:36:06 -04:00
Sébastien Han b70b098405 ceph: add support for stretched cluster crush rule
The pool spec has now a new property called "replicasPerFailureDomain"
which essentially represents the number of replicas to store in each
failure domain.
Assuming the failure domain is a datacenter (if the cluster is
stretched) then you will have 2 replicas per datacenter where each
replica ends up on a different host. This gives you a total of 4
replicas and for this, the "size" must be set to 4.

Closes: https://github.com/rook/rook/issues/5591
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-13 16:33:34 +02:00
Sébastien Han eba561f13e Merge pull request #6408 from leseb/fix-bz-1885971
ceph: reduce s3 max retry
2020-10-08 17:30:27 +02:00
Sébastien Han d01a35cadd ceph: reduce s3 max retry
When rgw is not responding the aws-sdk retries using exponential
backoff. Having 20 retries means we would wait for a while (hours) to
eventually fail. This is a problem as we cannot reflect the status of
the bucket properly.
Now we use 5 max retries which means around 8 seconds.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-08 16:20:00 +02:00
Sébastien Han 35d524c7ed Merge pull request #6322 from alimaredia/object-zone-unit-tests
ceph: add unit tests for object/zone
2020-10-08 10:24:05 +02:00
Ali Maredia b62a50309f ceph: add unit tests for object/zone
Add unit tests for CephObjectZone controller.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-10-07 18:52:44 -04:00
Travis Nielsen 3eddcf5477 Merge pull request #6283 from sp98/rook-IPv6-support
ceph: support IPv6 single-stack
2020-10-06 07:46:25 -06:00
Santosh Pillai 70c5f70701 ceph: support IPv6 single-stack for ceph
Add --ms-bind-ipv6 true args to mom, osd, mgr, object, mds, rbd daemons.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-10-02 01:03:03 +05:30
Travis Nielsen cebcf0a04e Merge pull request #6353 from leseb/fix-obc-external
ceph: fix obc upgrade from 1.3 to 1.4 external cluster
2020-10-01 07:37:35 -06:00
Sébastien Han d2d1da5970 ceph: fix obc upgrade from 1.3 to 1.4 external cluster
During the 1.3 cycle, we had implemented support for OBC external mode
through an Endpoint in the StorageClass parameter. In 1.4, this is not
the case anymore as we directly use the CephObjectStore external mode.
However, we must maintain backward compatibility with the cluster using
the Endpoint and not fail the upgrade.
Thus checking for a CephObjectStore is invalid and should not always be
assumed, now if we detect an Endpoint we return a simple context which
fixes the upgrade.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-01 11:33:54 +02:00
subhamkrai 0fddfcf307 ceph: handle golangci-lint linter errcheck error
this commit handle golangci-lint linter errcheck.

`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases

To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-30 22:24:34 +05:30
Jiffin Tony Thottan b7bc51ea69 ceph: remove user unlink while deleting OBC
User unlink is performed just before deleting OBC resources.
Since user is internal to OBC and next to remove to delete user,
unlinking user is not necessary.

Fixes: 6300
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-09-29 17:20:54 +05:30
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
Travis Nielsen bbb59d4949 Merge pull request #6274 from alimaredia/object-zonegroup-testing
ceph: add unit tests for zone group
2020-09-18 09:56:47 -06:00
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
Ali Maredia 2ffb4ab057 ceph: add unit tests for zone group
Add unit tests for CephObjectZoneGroup, with
minor fixes to the realm unit tests.

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-09-17 23:57:40 -04:00
subhamkrai f9fafe62d4 ceph: handle golangci-lint linter unused
this commit will enable one more linter
in golangci-lint.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 22:53:24 +05:30
Travis Nielsen 4df34f4bf4 Merge pull request #6226 from leseb/fix-6217
ceph: allow running rgw in rook with external mode
2020-09-14 11:18:53 -06:00