This field held state for a single reconciliation request, which
should not have been retrained / reused across multiple, possibly concurrent,
reconciliations.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.
Additionally, it silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
and string format errors found by golangci-lint, while at it.
Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
This reverts commit a941b3c33f.
Stop creating the 'cosi' user in the CephObjectStore reconcile. This
step often fails for some amount of time during initial object store
creation, causing frequent user concern. It has also been the source of
some reported failures that would otherwise be non-breaking for certain
users.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
For the specification see:
<https://github.com/rook/rook/blob/master/design/ceph/object/swift-and-keystone-integration.md>
* extend the API object specs for swift and keystone integration
* adapt rgw to the new go-ceph version
- The parameter lists of the API call have changes, as parameters
ignored by the RGW Admin Ops API are no longer serialized, therefore
the mock has to be adapted.
- There is now validation for the user keys that are passed to the
User get API, therefore things failed when we had empty keys in our
User proxy object.
* expand the reconcile loop for the swift and keystone integration
* fix minor mistakes in design document
* add env var to pass extra args to minikube
Minikube decides CPU cores and memory automatically based on the
available resources on the machine which may be insufficient to
run rook. This commit adds an environment variable to add arbitrary
arguments to the minikube command, so both can be specified if
desired.
* integration tests for swift and keystone
The new integration of swift or s3 and keystone support by rook
does not have any integration tests yet.
This commit introduces integration tests for swift and keystone. The
tests are done against a minimal keystone setup (keystone container
image from Yaook-project (https://yaook.cloud), sqlite as database
backend, cert-manager and trust-manager for test certificate setup).
To prevent hardcoded credentials, passwords are generated
by the tests. The integration tests use the openstack client
(keystone- and swift-functionality) (https://docs.openstack.org/
python-openstackclient/ latest/). This was a concious design decision
to use client tooling as close as possible to the end user instead of
using other go-libraries (such as gophercloud).
* add documentation on swift and keystone
Currently there is no documentation on the use of Swift to access
an object store as well as the use of OpenStack keystone for
authentication.
This commit adds documentation on the use of Swift and OpenStack
keystone, as well as CRD-related documentation and an example setup.
* add integration tests for S3 via keystone
This commit introduces integration tests for s3 and keystone. The
tests are run against the same minimal keystone setup that the tests
for swift and keystone use.
The integration tests use the aws s3 client to use client tooling as
close as possible to the end user instead of using other go-libraries.
Co-authored-by: Jan Klippel <jan.klippel@uhurutec.com>
Co-authored-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Signed-off-by: Sebastian Riese <sebastian.riese@cloudandheat.com>
Signed-off-by: Jan Klippel <jan.klippel@uhurutec.com>
Signed-off-by: Silvio Ankermann <silvio.ankermann@cloudandheat.com>
Add CephObjectStore spec.hosting.advertiseEndpoint configuration. This
provides a clear documented default for which endpoint Rook "advertises"
to dependent resources like CephObjectStores, OBCs, and COSI
Buckets/Accesses and allows users to override the default behavior if
desired.
The current default is to round-robin an endpoint from
spec.hosting.dnsNames, which has proven to be troublesome for some
users' object store configurations. This change provides much-needed
disambiguation for users.
This may be a breaking change for some existing spec.hosting.dnsNames
users. This is unexpected but is documented.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Loop variables cannot be reliably uses since they will
change with each iteration. Update these loop variable
uses to be safe by indexing the slice rather than
using the loop variable directly.
Also suppress the linter issues for passwords used
in tests.
Signed-off-by: travisn <tnielsen@redhat.com>
There are 2 cases of randomly generated secrets copied into Rook's unit
test code that have been flagged by Gitleaks. Add a comment to both
cases to help the tool understand that these aren't real production
secrets -- just unit test stand-ins.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The object user was previously required to be created in the
same namespace as the object store and the cluster. Now,
the object user can be reconciled even in a different namespace
from the cluster and object store. The namespace would be specified
in the object user CR.
Signed-off-by: travisn <tnielsen@redhat.com>
users & buckets admin capabilities aren't spelled the same
way between rook crd (singular) and ceph (plural). It's a bit
misleading when comparing ceph admin cap and rook users.
Signed-off-by: Peter Goron <peter.goron@gmail.com>
When creating a CephObjectStoreUser with a value spec.store that refers to an
unexisting CephObjectStore, after the reconciliation loop the
CephObjectStoreUser is in the ReconcileFailed state. However, a
ReconcileSucceeded event is created with this message:
"successfully configured CephObjectStoreUser"
The success message results of the return value for the error which is currently
`nil`. Let's replace it with the error message.
Signed-off-by: Lucas Henry <polyedre@disroot.org>
There is no reference for ssl in cephobjectstore Secret, so users won't
have much idea why tls secret need to used. Hence give reference
object stores tls secret ref in the Secret.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
The radosgw-admin command uses the network spec from ceph cluster spec
in object context but it is not filled properly in the object package.
But with PR 10898, network spec is available in clusterinfo which can
be used directly. Also removed cluserspec from object context.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Removing user caps will be skipped as the UserCaps is empty, which will cause updating user caps to not take effect.
Signed-off-by: Liang Zheng <zhengliang0901@gmail.com>
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor
Signed-off-by: parth-gr <paarora@redhat.com>
In object package code, external rgw server check is whether ceph cluster
is external instead of rgw. Correcting such scenarios in the code base.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
The CSI package needs to load clusterInfo, today this code is in the mon
package which makes the call of LoadClusterInfo impossible without
having a circular import.
Signed-off-by: Sébastien Han <seb@redhat.com>
The access/secrets keys for the user struct need to allocate only if
ceph user creation succeeds.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller
Closes: https://github.com/rook/rook/issues/9673
Signed-off-by: parth-gr <paarora@redhat.com>
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Since the mock client was added to the rgw/admin package from the
go-ceph project in https://github.com/ceph/go-ceph/pull/532 we can now
mock the HTTP client and validate the entire reconcile.
Signed-off-by: Sébastien Han <seb@redhat.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
The CephClusterSpec was missing from the object context, so the check
for the network provider in RunAdminCommandNoMultisite() was not
discovering the network property correctly.
Signed-off-by: Sébastien Han <seb@redhat.com>
The Rados Gateway Admin OPS API has changed its behavior from Nautilus
to Pacific. Calling user create on an existing user wil generate
additional keys to the user on Nautilus. Where in Pacific it will report
an error with UserAlreadyExists.
So to handle both scenarios, let's first get the user, and if the user
does not exist (NoSuchUser) we then create it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Prior to that the access and secret keys were left empty if the user
already existed, which led to updating the secret with empty values when
the operator restarts.
Signed-off-by: Sébastien Han <seb@redhat.com>
For the TLS communication for AdminOps Api, httpclient is required and
filled with TLS certs, currently it is set to nil pointer in
`buckethealthchecker` and `cephobjectstoreuser`.
Thanks @Krast76 finding the issue even proposing the fix.
Fixes: #8132
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.
Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
when store.Status was not initialized it caused
crashes when assigning store.Status.Info["endpoint"]
to objContext.Endpoint.
This commits makes sure store.Status is not Nil
before setting objContext.Endpoint and
objContext.SecureEndpoint in the ObjectStoreUser
controller.
Closes: https://github.com/rook/rook/issues/6916
Signed-off-by: Ali Maredia <amaredia@redhat.com>
This fixes libraries that were being imported multiple times in
pkg/operator/ceph/object and its subpackages.
Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
If the object store status is not yet initialized, the object store user should
fail its initialization and requeue the reconcile.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>