This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.
So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.
Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.
It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.
Signed-off-by: Sébastien Han <seb@redhat.com>
Adds rook-ceph-cluster chart to chart build, add tests.
Pulled out some common functionality between the Helm and non-Helm installers
Signed-off-by: Henry Zhang <me@henry.dev>
This commit updates external-snapshotter version to
v4.0.0 which supports snapshots v1.
Rook now defaults to enabling RBD and CephFS snapshotter
for K8s >= v1.17 and disabling it for K8s <= v1.16.
Supporting changes in documents and examples yaml files
are made.
Signed-off-by: Rakshith R <rar@redhat.com>
Rook's testing utilities include a function for loading the default
kubeconfig for a cluster. This function requires maintainance effort and
also doesn't support authentication methods like Basic Authentication.
Instead of adding support for it, drop the config loader and use the one
provided by the controller-runtime library.
Signed-off-by: Yannis Zarkadas <yanniszark@arrikto.com>
The integration tests run independently in the github actions so there is
no need to cleanup from every test. The cleanup is still needed in the
Jenkins tests where all the suites run serially.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When verifying OBC creation, validating that secret and configmap exist
is good, but the definitive validation is to check that the OBC's phase
is "Bound".
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The snapshot controller is specifying to use the canary image instead
of the expected version. Update the tests to set the correct version.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook's crashcollector pod posts entries to the ceph cluster when a crash occurs.
Over time the number cluster may hold crash entries needlessly.
To clean up old crash entries, this PR adds a field to the ceph cluster CR for the user to specify the number of days a crash entry should be kept for.
Providing a value for the field keepXDays creates a cronjob that runs every day at midnight, calling "ceph crash prune <keepXDays>".
Closes: https://github.com/rook/rook/issues/6332
Signed-off-by: Renan Campos <rcampos@redhat.com>
to run the integration test there needs
to be a few changes in manifests like using
`deviceFilter` and other related changes.
Signed-off-by: subhamkrai <srai@redhat.com>
This updates the chart to make use of helm3 which has been released
for some time, and also permits CRDs to be installed pror to other objects
allowing the chart to be deployed at the same time as CRs for rook objects.
Co-authored-by: Pete Birley <pete@port.direct>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
golangci-lint linter gosec showing more errors
than gosec gh. This commit resolve
new errors. And, removing
nosec comments from autogenerated files.
Signed-off-by: subhamkrai <srai@redhat.com>
this commit handle golangci-lint linter errcheck.
`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases
To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`
Signed-off-by: subhamkrai <srai@redhat.com>
this commit handle golangci-lint linter staticcheck error.
`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.
To see only `staticcheck` linter output
`golangci-lint run --disable-all -E staticcheck`
Signed-off-by: subhamkrai <srai@redhat.com>
this commit enables the rook integration
testing capability on openshift. currently
this commit enable CephSmokeSuite testing
only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
Added E2E testing to create,delete and restore
a snapshot, create a pvc-pvc clone, install
and uninstall snapshot controller and snapshot
CRD.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
This replaces any K8S clientset usages of the
`rbac.authorization.k8s.io/v1beta1` APIs with
`rbac.authorization.k8s.io/v1` client.
This is done as some RBAC objects used the `v1beta1` and others already
using `v1`, for consistency.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
`anon-user-access` clusterrolebinding is created when installing rook operator.
However each test suite does not delete the clusterrolebinding,
and cause errors.
This PR skip creating the clusterrolebinding in each suite if it already exists.
Signed-off-by: Takashi IIGUNI <iiguni.tks@gmail.com>
Currently, rook CI is not logging Kubernetes Events, but it sometimes gives hints
for bugs.
Therefore, this PR adds a function to collect and log Kubernetes Events.
Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.
A good status will look like:
status:
endpointStatus:
lastChanged: "2020-06-25T13:47:45Z"
lastChecked: "2020-06-25T13:48:46Z"
phase: Connected
A failed status:
status:
endpointStatus:
details: |-
error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
health: ERROR
This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.
Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
adds deploy.sh script to deploy validatingwebhookconfiguration and create secrets.
adds new command ceph admission-controller to start webhook servers.
adds validation for various rook custom resources
Signed-off-by: Vineet Badrinath <vbadrina@redhat.com>
There are several functions that are prefixed by "Is" and return
just error. It's straightfoward to return bool to make the meaning
of these functions clearer.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The CRD watcher has been replaced by the new controller-runtime
framework.
This brings robustness in our operator, meaning that any resources that
are modified will be reconciled into the desired state.
Closes: https://github.com/rook/rook/issues/4940
Signed-off-by: Sébastien Han <seb@redhat.com>
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The methods and arguments to the exec methods are not all used anymore.
This cleans up the methods to only what is necessary to improve
the readability and maintainability.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests have failed intermittently due to needing just
a little longer to start the file test pod. This increases the
wait timeout.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests have been mostly running on the flex driver
with only a newer test on the csi driver. With the CSI driver being
the preferred driver going forward, now the integration tests will
all be running with the CSI driver with the exception of a test
suite that is only dedicated to the flex driver.
A number of other test improvements are also made for code
readability, test stability, and removing unused options.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When the tests run in a PR, they can only run against a single version
of K8s by default. If more than five k8s versions are supported in hte
CI, we will need to run some of the suites on multiple versions.
The versions are comma-separated in the list.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Sometimes we fail to get logs for pods using this method, notably the
operator pod. It is unknown why this happens. Pod logs are VERY
important, so try again using kubectl.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
Reset the integration test k8s helper's `RetryLoop` to its original
value, and instead only wait an extra long time to allow the mgr module
updates to take a long time after Ceph is updated from Mimic to Nautilus
as part of Ceph's upgrade integration test.
Updating mgr modules can hang for quite a while, which causes the tests
to time out waiting for the OSDs to be updated. Allow this to take a
long time so the tests aren't as flaky.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
During upgrade tests, Rook should verify that it can still run legacy
OSDs. This includes directory-based OSDs, filestore disk OSDs, and
bluestore disk OSDs installed without ceph-volume (i.e., before mimic
v13.2.2) can still be run after upgrade.
This necessitates running the upgrade test twice; once with filestore
and once with bluestore.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>