This reverts commit a941b3c33f.
Stop creating the 'cosi' user in the CephObjectStore reconcile. This
step often fails for some amount of time during initial object store
creation, causing frequent user concern. It has also been the source of
some reported failures that would otherwise be non-breaking for certain
users.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
Given that Ceph Quincy (v17) is past end of life,
remove Quincy from the supported Ceph versions,
examples, and documentation.
Supported versions now include only Reef and Squid.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For image mode mirroring, if cephBlockPool.Pool.Spec.Mirroring.Enable
is set to false, then remove the peer cluster and disable mirroring on
all the pool if the user has disabled mirroring on all the pool images.
If mirroring is not disabled on all the pool images, then reconcile will
fail asking the users to manually disable mirroring on those images.
Signed-off-by: sp98 <sapillai@redhat.com>
Rook has been setting the application automatically on all
pools to rbd for CephBlockPools, rook-ceph-rgw for
CephObjectStores, mgr on the built-in .mgr pool,
and nfs on the built-in .nfs pool.
The legacy pool device_health_metrics is long gone
from Pacific which is no longer supported, so we can
remove special handling for that pool in the upgrade
guide and in the code.
The application setting is now available on the pool spec
although it is not expected to commonly need to override
the default applications set by Rook.
The application for CephFilesystem pools is now being
set to cephfs, where it was previously blank.
Signed-off-by: travisn <tnielsen@redhat.com>
Loop variables cannot be reliably uses since they will
change with each iteration. Update these loop variable
uses to be safe by indexing the slice rather than
using the loop variable directly.
Also suppress the linter issues for passwords used
in tests.
Signed-off-by: travisn <tnielsen@redhat.com>
In Reef the is_master changed from a string to a bool
so we must update the type for proper json
serialization.
Signed-off-by: travisn <tnielsen@redhat.com>
There are 2 cases of randomly generated secrets copied into Rook's unit
test code that have been flagged by Gitleaks. Add a comment to both
cases to help the tool understand that these aren't real production
secrets -- just unit test stand-ins.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The bucket health checker was removed in 1.10. Now in 1.12
we no longer need this removal of the bucket health
checker since it will no longer exist to remove.
Signed-off-by: travisn <tnielsen@redhat.com>
External CephObjectStores already have endpoints defined by
spec.gateway.externalRgwEndpoints, and if the external store is
configured with TLS (HTTPS), the store's certificates will likely not
accept connections intended for the Service endpoint Rook creates. Some
users might not be able to easily add the service endpoint to their
certificates. Therefore, don't even bother creating a Service for
external clusters.
This does introduce a few issues. The Service seems to have been
initially created to allow multiple external RGW endpoints to be
addressable via a single address in Rook. For all connections to an
external CephObjectStore with multiple endpoints, simply choose an
endpoint at random. Random selection will prevent Rook from failing to
create buckets or users on an external store if one of the external
store's endpoints fails.
The latest OBC library (lib-bucket-provisioner) allows updating the
endpoints on ObjectBuckets after they are created. This allows Rook
users to change endpoints on external CephObjectStores without breaking
all existing OBCs. It requires implementation of the new GetUserID()
library call, requires updating Provision() and Grant() calls to be
idempotent, and it requires removing the Update() call.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Remove the health checker for CephObjectStore. The liveness and
readiness probes go through the same code paths in RGW as creating
buckets without as much affect on the storage backend.
Full discussion: https://github.com/rook/rook/issues/11031
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
User can define his desired endpoint list in Zone CR so that it will
overwrite the default service name for rgw.
Resolves#6432
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Check the radosgw-admin realm user list per object store instead of relying
on the ceph dashboard get-rgw-api- command.
Resolves#9099
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
This removes double package imports. Example:
```
"github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
cephv1 "github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
```
Only one is now being used as shown in go-staticcheck ST1019
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
Added few test cases to check for the following:
* creation of an external object store
* deletion of an external object store
* creation of an external object store with a missing secret
* creation of an external object store with ExternalRGWEndpoints set to
nil.
Signed-off-by: Pranshu Srivastava <rexagod@gmail.com>
We need to explicitly assign the ceph version within each scope since
one function returns a pointer and one is not.
Previously, the `desiredCephVersion` was recreated in the scope of the
`else` and thus a new variable `desiredCephVersion` was initialized (see
the `:=`).
Now, each scope creates and assign to clusterInfo a `desiredCephVersion`
value.
Signed-off-by: Sébastien Han <seb@redhat.com>
The reconcile was skipping updating most pool properties for
object stores. The implementation of pools between the file,
object, and pool controllers had some duplicate code, so
this change also factors out the common code for better
reuse in a single place. Anytime a pool is created or updated,
it will now consistently update all the pool properties
that are expected to be modifiable.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Prior to this, we were comparing a pointer (the memory address) with a
struct. This was obviously always failing and returned false. We must
dereference the pointer to access the data contained at that memory
location.
Closes: https://github.com/rook/rook/issues/9544
Signed-off-by: Sébastien Han <seb@redhat.com>
The original PR which added event reporting unnecessarily "optimized" to
prevent spamming the API controller[1].
It is sometimes important to get events as they happen and not hide new
events behind preexisting older events. For example, in integration
tests, we may often want to wait for a controller to finish processing
an update, and the best way to do that is to wait for the
"ReconcileSucceeded" event. In order for this to be useful, the events
must be reported each time.
If we begin having problems with events being reported too often, then
we should fix the underlying issue of reconciles happening too often
instead of relying on a time-based "optimization" that hides recent
event reports that may be useful.
[1]: https://github.com/rook/rook/pull/7222
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
Add to the RGW multisite integration test a verification that the RGW
period is committed on the first reconcile and not committed on the
second reconcile.
Do this in the multisite test so that we verify that this works for
both the primary and secondary multi-site cluster.
To add this test, the github-action-helper.sh script had to be modified
to
1. actually deploy the version of Rook under test
2. adjust how functions are called to not lose the `-e` in a subshell
3. fix wait_for_prepare_pod helper that had a failure in the middle
of its operation that didn't cause failures in the past
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Replace calls to 'radosgw-admin period update --commit' with an
idempotent function.
Resolves#8879
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
If the CephObjectStore health checker fails to be created, return a
reconcile failure so that the reconcile will be run again and Rook will
retry creating the health checker. This also means that Rook will not
list the CephObjectStore as ready if the health checker can't be
started.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.
Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
When creating object store, `radosgw-admin realm get ..` command is stuck forever when required number of OSDs are not available. Because of this the uninstall of object store is also stuck. User has to manually remove the finalizer to delete the object store.This PR uses `ExecuteCommandWithTimeout` for running `radosgw-admin` command. Timeout during installation will be reconciled. Cleanup will be treated as best effort. Any errors during uninstalling of single site object store will only be logged.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Add unit tests for when the "Zone" is configured
in an object store so that the object store joins
the CephObjectZone and the zone's corresponding
multisite configuration.
Signed-off-by: Ali Maredia <amaredia@redhat.com>
this commit handle golangci-lint linter staticcheck error.
`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.
To see only `staticcheck` linter output
`golangci-lint run --disable-all -E staticcheck`
Signed-off-by: subhamkrai <srai@redhat.com>
this commit will enable one more linter ineffassign
in golangci-lint.
This linter throws an error when variable is assigned and never used.
`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
By addding ".svc.cluster.local" to the endpoint of the object store, the
OBC requests won't go through a proxy if any is configured.
Signed-off-by: Sébastien Han <seb@redhat.com>
Instead of setting the Phase of the CR once the reconcile is done, let's
actually set in from the healthcheck so it is more accurate.
Closes: https://github.com/rook/rook/issues/5249
Signed-off-by: Sébastien Han <seb@redhat.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Now, when an S3 user gets created, Rook will add the S3 endpoint to the
Secret along with the credentials.
Signed-off-by: Sébastien Han <seb@redhat.com>
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.
A good status will look like:
status:
endpointStatus:
lastChanged: "2020-06-25T13:47:45Z"
lastChecked: "2020-06-25T13:48:46Z"
phase: Connected
A failed status:
status:
endpointStatus:
details: |-
error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
health: ERROR
This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.
Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
Let's not wait for the CephCluster to be done reconciling but instead
check for the Ceph cluster status, if it's closed to "ok" then we
proceed so HEALTH_OK and HEALTH_WARN are accepted.
Signed-off-by: Sébastien Han <seb@redhat.com>
This is critical to get this information so that the label can be set
properly and also the configuration done appropriately.
Signed-off-by: Sébastien Han <seb@redhat.com>