Commit Graph
64 Commits
Author SHA1 Message Date
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Sébastien Han 86c9a8f3de rgw: read tls secret hint for insecure tls
If the admin wants to use insecure TLS to validate connections to rgw
internally, the TLS secret can have another entry "insecureSkipVerify"
and set it to "true".

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-28 11:27:00 +02:00
Sébastien Han cda5dad291 rgw: use insecure TLS for bucket health check
We have seen cases where the signed certificate used for the RGW does not
contain the internal DNS endpoint, resulting in the health check to fail
since the certificate is not valid for this domain.
People consuming the gateways by external clients and for specific
domains do not necessarily have the internal DNS configured in the
certificate.
So let's be a bit more flexible and simply ensure a connectivity check
and bypass the certificate validation.

Also, this is fixing the tls code in newS3Agent and adds unit tests.

Closes: #8663
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-28 14:48:07 +02:00
Jiffin Tony Thottan 280c29f330 ceph: pass region to newS3agent()
If the region is specified in the storage class of OBC, use that in the
newS3agent() than using constant "us-east-1".

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-09-21 13:08:18 +05:30
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Jiffin Tony Thottan f4bb47e440 ceph: add support for update() from lib-bucket-provisioner
Recently lib-bucket-provisioner add support for update() API.
Include that on the obc implementation since it can be used to
update quota for OBC.

Fixes: #7146

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-08-19 23:57:11 +05:30
Sébastien Han ad45924807 ceph: change the debug implementation of the admin ops API
The go-ceph library has removed the `Debug` field from the API type in
https://github.com/ceph/go-ceph/pull/543. Since the HTTP Client can be
mutated we now have our own client to dump requests and responses when
the operator log level is DEBUG.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-29 13:59:50 +02:00
Sébastien Han 1fd76a1168 ceph: add missing ceph cluster spec
The lib-bucket-provisioner was missing the cephCluster spec so the check
in RunAdminCommandNoMultisite would fail.
This is similar to ad54ed8aac just a
different place.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-13 18:45:39 +02:00
Sébastien Han f074c12c4d ceph: get the s3 user first instead of create
The Rados Gateway Admin OPS API has changed its behavior from Nautilus
to Pacific. Calling user create on an existing user wil generate
additional keys to the user on Nautilus. Where in Pacific it will report
an error with UserAlreadyExists.
So to handle both scenarios, let's first get the user, and if the user
does not exist (NoSuchUser) we then create it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-09 18:32:48 +02:00
Sébastien Han 97f1cc3fd2 Merge pull request #8244 from thotz/obcfetchcorrectport
ceph: fetch rgw port from cephobjectstore crd for obc
2021-07-05 11:21:48 +02:00
Jiffin Tony Thottan e4a9b4bac1 ceph: fetch rgw port from cephobjectstore crd for obc
Currently `setObjectStorePort()` queries port from RGW service and
picks up the first port from the list, other places of the code
prefers for TLS, if it is enabled so for OBC's we end up having
`https` endpoint with non-secure port.

Fixes: #8141
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-07-05 10:55:05 +05:30
Jiffin Tony Thottan 560f44b07e ceph: enable debug for adminops client if the rook loglevel <= debug
The debug for adminOps client can be enabled by setting `Debug` flag.
In this PR, it is enabled for OBC and cephobjectstore not healthchecker.
Also add similar changes while calling `NewS3Agent` in the bucket
provisioner code path.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-07-01 11:36:51 +05:30
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Jiffin Tony Thottan 867474e405 ceph: initialise httpclient for bucketchecker and objectstoreuser
For the TLS communication for AdminOps Api, httpclient is required and
filled with TLS certs, currently it is set to nil pointer in
`buckethealthchecker` and `cephobjectstoreuser`.

Thanks @Krast76 finding the issue even proposing the fix.

Fixes: #8132
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-06-21 10:59:58 +05:30
Sébastien Han 90bea8a560 ceph: stop using radosgw-admin CLI for s3 user management
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.

Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 11:08:22 +02:00
Jiffin Tony Thottan f9c7634d88 ceph: support for OBC operations on enabled TLS RGW endpoint
OBC operations will fail if RGW only enabled HTTPS. With this change, OBC operations
will use needed TLS certs to succeed the operation.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-28 12:37:38 +05:30
Jiffin Tony Thottan bcbe707bcb ceph: fix healthcheck incase ssl enabled for rgw
In case ssl enabled for RGW, bucket healthcheck won't work since access is denied from the server.
In that case configure S3 agent using "InsecureSkipVerify: true" option.

Fixes: 7288
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-04-20 14:23:48 +05:30
Blaine Gardner b8dc2a3214 ceph: update lib bucket provisioner
Update to the latest lib bucket provisioner code.
Fixes issue 6650

Modifies CRD for objectbucketclaims to fix an additional bug where an
ObjectBucket's 'ClaimRef' is lost due to the CRD validation being
specified incorrectly.

Changes OBC deletion/cleanup to delete the bucket before the user. A
user cannot be deleted without an unsafe purge option if the user has
buckets associated to it.

Does not reintroduce bug 6767 from previous fix for 6650

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-18 12:48:54 -07:00
Blaine Gardner bc08f51f65 Merge pull request #6699 from BlaineEXE/update-lib-bucket-provisioner
ceph: update object bucket provisioner library
2020-12-03 14:47:44 -07:00
Blaine Gardner 3b4ee6c8e3 ceph: update object bucket provisioner library
The library for object bucket provisioning is updated to fix errors
during provisioning.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-03 09:43:12 -07:00
Blaine Gardner e75dcac8a3 ceph: add more debug logging to obj bucket claims
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-01 14:01:33 -07:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Travis Nielsen cebcf0a04e Merge pull request #6353 from leseb/fix-obc-external
ceph: fix obc upgrade from 1.3 to 1.4 external cluster
2020-10-01 07:37:35 -06:00
Sébastien Han d2d1da5970 ceph: fix obc upgrade from 1.3 to 1.4 external cluster
During the 1.3 cycle, we had implemented support for OBC external mode
through an Endpoint in the StorageClass parameter. In 1.4, this is not
the case anymore as we directly use the CephObjectStore external mode.
However, we must maintain backward compatibility with the cluster using
the Endpoint and not fail the upgrade.
Thus checking for a CephObjectStore is invalid and should not always be
assumed, now if we detect an Endpoint we return a simple context which
fixes the upgrade.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-01 11:33:54 +02:00
Jiffin Tony Thottan b7bc51ea69 ceph: remove user unlink while deleting OBC
User unlink is performed just before deleting OBC resources.
Since user is internal to OBC and next to remove to delete user,
unlinking user is not necessary.

Fixes: 6300
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-09-29 17:20:54 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
subhamkrai f9fafe62d4 ceph: handle golangci-lint linter unused
this commit will enable one more linter
in golangci-lint.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 22:53:24 +05:30
subhamkrai 209dc6251d ceph: fixed DNS suffix breaks in custom DNS suffix clusters
By addding ".svc" to the endpoint of the object store,
will fix the custom DNS suffix cluster.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-11 13:29:53 +05:30
Sébastien Han 96b4a56a57 ceph: silence aws s3 sdk logs on healthcheck
We don't need to activate the debug logs on the objectstore
healthchecks.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-01 09:32:15 +02:00
Sébastien Han 3a92f98233 ceph: use special dns suffix for object ep
By addding ".svc.cluster.local" to the endpoint of the object store, the
OBC requests won't go through a proxy if any is configured.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-08-26 17:10:31 -06:00
Jiffin Tony Thottan 7413485b8a ceph: use proper log functions in object package
There is incorrect usage of log functions in the object directory, changing those instances

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-08-21 12:39:53 +05:30
Jiffin Tony Thottan 091679cbeb ceph: use deleteOBCResourceLogError instead of deleteOBCResource
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-07-31 15:41:35 +05:30
Jiffin Tony Thottan 59f1e8364e ceph: fix up deleteOBCResourceLogError
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-07-31 15:41:31 +05:30
Travis Nielsen e0ea87a353 Merge pull request #5674 from thotz/quotaobc
ceph: adding support for quota in object bucket claims
2020-07-30 22:51:58 -06:00
subhamkrai 0279025e9e ceph: handling all the gosec errors
a few of the gosec errors were left. so
this commit will resolve all the errors.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-30 08:19:55 +05:30
Jiffin Tony Thottan f4240cb31d ceph: add quota support for obc
Closes: #5274
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-07-28 12:00:11 +05:30
Jiffin Tony Thottan cf61966205 ceph: remove rgwerrno from setuserquota
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-07-28 11:59:51 +05:30
subhamkrai acc4ed5df9 ceph: handling gosec errors that are not checked
this commit handles all the gosec g104
i.e audit errors not checked inside pkg.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-24 20:02:06 +05:30
Ali Maredia e6ed4ff8ea ceph: minor fixes + add realm/zg/zone to object context
- Add realm/zone group/zone to Object Context,
so that any call to runAdminCommand has the
same realm, zone group, and zone as the object
store that the call is going to.

- Also modify the delete code for the object-store
using multisite to remove the endpoints of an
object store from a zone instead of the normal
clean-up

- Added more debug logging all around the object-store
code related to multisite

- files generated by rerun of `make codegen`

- change back the edit on the Copyright in object.go

Signed-off-by: Ali Maredia <amaredia@redhat.com>
2020-07-21 16:47:59 -04:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Travis Nielsen 4b45048557 ceph: skip logging output of bucket policy change
The output of the bucket policy change should not be logged in case
there is sensitive information in the output.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-14 11:50:16 -06:00
Sébastien Han e4eaa91ede ceph: add rgw endpoint healthcheck
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.

A good status will look like:

status:
  endpointStatus:
    lastChanged: "2020-06-25T13:47:45Z"
    lastChecked: "2020-06-25T13:48:46Z"
  phase: Connected

A failed status:

status:
  endpointStatus:
    details: |-
      error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
      caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
    health: ERROR

This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.

Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-02 16:34:10 +02:00
Sébastien Han 5f74e493ef ceph: small user delete refactor
Do not return error code, interpret it directly, put the error as part
of the output on failures.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-18 18:49:03 +02:00
Jiffin Tony Thottan 96e3d1529e ceph: remove unlinking user in revoke api for obc
UnlinkUser works if bucket owner is obc user, otherwise it result in mismatch.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-06-10 18:30:21 +05:30
Jiffin Tony Thottan c0e23e0e16 ceph: clean up resources incase of failure
When obc creation fails in Grant() or Provision() the resources such as rgw user or bucket
created while provisioning need to remove.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-06-10 18:30:16 +05:30
Jiffin Tony Thottan 07f5aa7a19 ceph: remove deleteCephUser() and deleteBucket()
Replacing deleteCephUser() and deleteBucket() with deleteOBCResource()
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-06-10 18:30:12 +05:30
Jiffin Tony Thottan 9fccfb0e6f ceph: introduce new function for cleaning resources
Combining DeleteUser() and DeleteBucket() into single function
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2020-06-10 18:30:06 +05:30
Sébastien Han fc648970c4 Merge pull request #5511 from thotz/existingbucketsuserpolicy
obc: Do not delete user for retain policy if it is owner
2020-06-03 14:11:57 +02:00
Sébastien Han a90c18e0cb ceph: fix obc additionalConfig field
Some external consumers of Rook-Ceph are using the
lib-bucket-provisioner operator which creates different CRDs for both
objectbuckets.objectbucket.io and objectbucketclaims.objectbucket.io.
They have the `additionalConfig` declared in their spec. Rook-Ceph does
not and thus ignores it.
For the OB creation to work, the `additionalConfig` must be acknowledged
by the operator and populate or initialize it if empty.

This is preventing the following error:

controller.go:190] error syncing 'default/ceph-retain-bucket-converged': error creating OB "": ObjectBucket.objectbucket.io "obc-default-ceph-retain-bucket-converged" is invalid: spec.endpoint.additionalConfig: Invalid value: "null": spec.endpoint.additionalConfig in body must be of type object: "null", requeuing

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-02 16:04:20 +02:00