The generation of a long node name in the integration tests was
being done based on the k8s version. In the past, older K8s versions
did not support the changing name. Now it's more maintainable if
we generate the long name depending on the test suite.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The upgrade integration test was from rook v1.6 to the latest master.
This was necessary until we are ready for the v1.8 release, from which
time we want to focus the upgrade testing from v1.7 to the latest
master.
The duplication in the test CRs and other resources is now reduced
by the upgrade calling a thin wrapper to forward a call to the
master version of the resource. When a new feature is added that
needs to be differentiated from the previous version, the method
then can be implemented instead of wrapping the master implementation.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ingress api version changed when it went to v1, and this has caused some upheaval
throughout the kubernetes ecosystem. This commit uses a common method of deciding which
ingress api to use, and allows the optional override of the kubernetes version
presented to helm using the helm build-in capabilities.
also add an ingress into the helm integration tests so any regressions to how ingresses
are handled in the future are caught easier.
Closes rook#9174
Signed-off-by: Tom Hellier <me@tomhellier.com>
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to k8sutil node functions. By this,
we can handle cancellation during API call of node resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil pod functions. By this, we
can handle cancellation during API call of pod resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.
Closes: #8407
Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The volume replication CRDs are an external component, not owned by Rook.
Therefore, they should be installed as any other independent component
in case the admin will install other consumers of the volumereplication CRDs
in the future in addition to Rook and the CSI driver.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We have seen cases where the signed certificate used for the RGW does not
contain the internal DNS endpoint, resulting in the health check to fail
since the certificate is not valid for this domain.
People consuming the gateways by external clients and for specific
domains do not necessarily have the internal DNS configured in the
certificate.
So let's be a bit more flexible and simply ensure a connectivity check
and bypass the certificate validation.
Also, this is fixing the tls code in newS3Agent and adds unit tests.
Closes: #8663
Signed-off-by: Sébastien Han <seb@redhat.com>
The `mgr host ls` looks like:
`"addr": "10.1.0.27/fv-az244-362", "hostname": "fv-az244-362"`
Previously, we were picking `addr` which contains both IP and hostname
where the k8snode has the hostname only. So let's use the Hostname from
the mgr command output instead.
Signed-off-by: Sébastien Han <seb@redhat.com>
In general, make the output from the object store test easier to follow
and debug.
- Make sure assertions are associated with the appropriate subtest.
- Remove `t.Helper()` from complex helpers.
- Add info to log messages.
- Move long parts of (sub-)test names to log messages.
- Move some reused functions back to ceph_base_object_test.go
- Refactor runObjectE2ETestLite() to be suitable for use as the second
store created in the "run a second object store" test.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
We have new jobs now:
* one that runs both smoke and object on the next Pacific version
* one that runs both smoke and object on Ceph master
* one that tests the upgrade from the current pacific stable to the
pacific devel
* one that tests the upgrade from the current octopus stable to the
octopus devel
Signed-off-by: Sébastien Han <seb@redhat.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Multiple CephObjectStores should be able to run at the same time. In CI,
test this by:
1. Create the primary object store under test (store A) (this was
already done)
2. Verify that store A starts and becomes ready (already done but is
moved here)
3. Create a second object store (store B) (new)
4. Verify that store B starts and becomes ready (new)
5. Delete store B, and ensure it is deleted (new)
6., etc. Continue to verify primary store A operations as before (this
was already done)
This is necessary to ensure that the Rook operator is using
`radosgw-admin` and the RGW Admin Ops API properly to support multiple
ObjectStores running at once.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The integration tests must always be run against the local
build of rook, and an image should never be pulled from dockerhub.
To prevent pulling a release or master tag, the local build
will use a tag specific to the build and not ever published
elsewhere.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit a8a40428b0)
The MDS core team suggested with deploy the MDS daemon first and then do
the filesystem creation and configuration. Reversing the sequence lets
us avoid spurious FS_DOWN warnings when creating the filesystem.
Closes: #8745
Signed-off-by: Sébastien Han <seb@redhat.com>
If the region is specified in the storage class of OBC, use that in the
newS3agent() than using constant "us-east-1".
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
The `*SuiteMinimalTestVersion` vars in
`tests/integration/ceph_base_deploy_test.go` are no longer needed with
Jenkins no longer being used. Remove them.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The CephSmokeSuite is becoming quite large and long, and most of the
length is now related to the object e2e tests. Separate the full object
e2e test into CephObjectSuite, and only test the object 'lite' test in
the CephSmokeSuite.
Resolves#8714
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
We just need to wait longer when the status is not ready. We needed
another sleep otherwise the status was never nil and the loop went too
fast. See:
```
2021-09-02 16:29:11.372249 I | integrationTest:
2021-09-02 16:29:11.374427 I | integrationTest:
2021-09-02 16:29:11.377764 I | integrationTest:
2021-09-02 16:29:11.379950 I | integrationTest:
2021-09-02 16:29:11.382084 I | integrationTest:
2021-09-02 16:29:11.385383 I | integrationTest:
2021-09-02 16:29:11.388499 I | integrationTest:
2021-09-02 16:29:11.391301 I | integrationTest:
2021-09-02 16:29:11.393545 I | integrationTest:
2021-09-02 16:29:11.396249 I | integrationTest:
```
Signed-off-by: Sébastien Han <seb@redhat.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
The Cassandra tests are failing in the CI, although they pass when
run locally on a developer cluster. Until the issue can be tracked
down we need to disable these checks to get to a green CI again.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
we are adding new linter for shellcheck.
As we are writing more shell scripts this
will help maintain quality.
Also, doing all the changes required in
bash files to pass this shellcheck.
Closes: https://github.com/rook/rook/issues/8431
Signed-off-by: subhamkrai <srai@redhat.com>
Recently lib-bucket-provisioner add support for update() API.
Include that on the obc implementation since it can be used to
update quota for OBC.
Fixes: #7146
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
Rook orchestrator is going to use PVs as <devices> for OSDs.
These PVs must use one storage class present in the cluster and
configured properly in the manager.
With this change we include this new requirement in the Rook manager module test
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
Enable again the mgr test:
- Now is more reliable and robust the start of the test.
- Minor fixes to adapt the <service ls> to the new name of the crash daemon
deployed by rook
- The creation of OSDs is disabled in the orchestrator, so i have removed
this test until we will have the functionality ready again in the orchestrator
part
My plan is to provide in the orchestrator two different ways to create OSDs:
- Creation of OSD using specific devices if discovery daemon is running
- Creation of OSds using PVS (if we have LSO/other LS operator running)
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
With v1.7 approaching, the upgrade integration test will now test from
v1.6.x to the latest master, which will effectively become the v1.7
release soon.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.
So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.
Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.
It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.
Signed-off-by: Sébastien Han <seb@redhat.com>
The helm tests frequently hit an issue with the admission controller.
The admission controller is also enabled in other test suites, so
it seems safe enough to disable for the helm test suite.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.
So the automatic configuration of Ceph Filesystem peers is now possible.
By editing the CephFilesystem CRD, you can now turn on mirroring:
```yaml
mirroring:
enabled: false
# list of Kubernetes Secrets containing the peer token
# for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
peers:
secretNames:
- secondary-cluster-peer
```
Also, the mirroring status is displayed in the CR status:
```
status:
info:
fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
mirroringStatus:
daemonsStatus:
- daemon_id: 4186
filesystems:
- filesystem_id: 2
name: myfs
lastChecked: "2021-07-01T14:16:29Z"
phase: Ready
snapshotScheduleStatus:
lastChecked: "2021-07-01T14:16:29Z"
snapshotSchedules:
- fs: myfs
path: /
rel_path: /
retention: {}
schedule: 24h
```
Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>