Commit Graph
523 Commits
Author SHA1 Message Date
subhamkrai f17f908c31 ceph: disable admission controller for CephUpgradeSuite
disabling admission controller for the upgrade test.
Upgrade test was using older controller runtime version
and api v1 need latest version controller runtime(v0.7).

Signed-off-by: subhamkrai <srai@redhat.com>
2021-01-21 22:26:02 +05:30
Travis Nielsen 1e111897a9 ceph: allow progressing status during integration test cleanup
The integration tests check after every test suite that the cluster is
either ready or connected. Sometimes the operator is in progressing state
because the reconcile may be triggered by some event that is unpredictable
at the end of the test. For test stability we allow the progressing status.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-13 15:58:35 -07:00
Travis Nielsen 08c227dbdd ceph: remove unexpected array type in crd schema for tests
Clean up an unexpected array type in the test schema for pre-k8s-1.16
clusters

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-13 11:20:53 -07:00
Travis Nielsen 1705c64ced ceph: add devices to schema at overall storage level
With the 1.5 schema added to the CRDs, the devices at the root
level of the storage element were missed. Now the ability to
specify devices at the root storage level to apply to all
nodes is restored.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-11 11:46:27 -07:00
Blaine Gardner 9b0ba6ae8b ceph: add obc to upgrade test
Add object bucket claim to Ceph upgrade test.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-08 09:56:00 -07:00
Blaine Gardner 40fd80cf14 ceph: update smoke test to verify obc is bound
When verifying OBC creation, validating that secret and configmap exist
is good, but the definitive validation is to check that the OBC's phase
is "Bound".

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-01-07 13:58:22 -07:00
Blaine Gardner b8dc2a3214 ceph: update lib bucket provisioner
Update to the latest lib bucket provisioner code.
Fixes issue 6650

Modifies CRD for objectbucketclaims to fix an additional bug where an
ObjectBucket's 'ClaimRef' is lost due to the CRD validation being
specified incorrectly.

Changes OBC deletion/cleanup to delete the bucket before the user. A
user cannot be deleted without an unsafe purge option if the user has
buckets associated to it.

Does not reintroduce bug 6767 from previous fix for 6650

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-18 12:48:54 -07:00
Travis Nielsen 114a949cfb ceph: test latest ceph on raw devices and previous ceph on partitions
The latest ceph-volume does not support creating OSDs on partitions.
The github actions only have a partition available, so we will run
the github tests on v14.2.12 and v15.2.7 that still support partitions,
while the Jenkins environment has a raw device available where we
can run the latest versions of Ceph in the tests.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-17 18:21:16 -07:00
Travis Nielsen 2f59a96bcf Merge pull request #6840 from travisn/snapshot-test-version
ceph: use the correct snapshot controller version in the tests
2020-12-16 11:40:17 -07:00
Travis Nielsen c15499ed7e ceph: use the correct snapshot controller version in the tests
The snapshot controller is specifying to use the canary image instead
of the expected version. Update the tests to set the correct version.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-16 10:25:00 -07:00
Travis Nielsen 55203f5069 build: replace hostpath provisioner with static pvs
The hostpath provisioner has never been supported and neither is it
working on K8s 1.20. For the tests we will simply use static
local PVs so we can avoid the whole problem of an unsupported
hostpath provisioner.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-15 22:13:25 -07:00
Blaine Gardner 26c9512962 ceph: add namespace meta-comments to manifests
Add meta-comments to manifests to allow basic templating via `sed`.

Add the following types of meta-comments:
- # namespace:X
  - A basic namespace
  - replace the field with a namespace for X
- # serviceaccount:namespace:X
  - A service account namespace for SCC
  - e.g., "system:serviceaccount:<ns>:rook-ceph-system"
  - Replace the namespace "<ns>" with with a namespace for X
- # provisioner:namespace:X
  - A provisioner identifier with namespace prefix
  - e.g., "<ns>.cephfs.csi.ceph.com"
  - e.g., "<ns>.ceph.rook.io/bucket"
  - Replace the namespace "<ns>" with a namespace for X

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-09 16:21:31 -07:00
Sébastien Han 8aaff235bb ceph: ability to set external services ports
When the cluster is external we want to expose manager service port
along with the endpoints.
This allows us to connect but Rook will keep on using 9283 as a facing
port for Prometheus and more.

```
monitoring:
  enabled: true
  ...
  ...
  externalMgrPrometheusPort: 9283
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-12-04 09:52:09 +01:00
subhamkrai c22590e2cb ci: helm integration test is failing on pre-k8s-1.16
testCephHelmSuite is failing in pre-k8s v1.16
due to invalid schema. v1beta1 doesn't support
null value so returning quotes instead of an empty
string.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-12-04 10:52:32 +05:30
Blaine Gardner 3b4ee6c8e3 ceph: update object bucket provisioner library
The library for object bucket provisioning is updated to fix errors
during provisioning.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-12-03 09:43:12 -07:00
Blaine Gardner 87ee8dce29 Merge pull request #6679 from leseb/add-log-collector
ceph: add log collector
2020-12-01 12:08:50 -07:00
Sébastien Han c6a87203ca ceph: add log collector
We can now collect logs directly into a side-car container.
A new CRD spec has been added:

spec:
  logCollector:
    enabled: true
    periodicity: 24h

Every 24h we will rotate log files for each Ceph daemon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-12-01 16:30:17 +01:00
Satoru Takeuchi a17cec45a6 Merge pull request #6709 from leseb/fix-6708
ceph: fix RBAC for cron cash pruner
2020-11-27 15:18:36 +09:00
Sébastien Han 13da4a1521 ceph: fix RBAC for cron cash pruner
The rook-ceph-system service account was lacking delete permission and
the operator was complaining.

Closes: https://github.com/rook/rook/issues/6708
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-26 11:29:47 +01:00
Renan Campos ad9dcb0c05 ceph: periodically prune crash entries older than user-provided days
Rook's crashcollector pod posts entries to the ceph cluster when a crash occurs.
Over time the number cluster may hold crash entries needlessly.
To clean up old crash entries, this PR adds a field to the ceph cluster CR for the user to specify the number of days a crash entry should be kept for.
Providing a value for the field keepXDays creates a cronjob that runs every day at midnight, calling "ceph crash prune <keepXDays>".

Closes: https://github.com/rook/rook/issues/6332
Signed-off-by: Renan Campos <rcampos@redhat.com>
2020-11-24 17:06:18 -05:00
Sébastien Han f86f8300f9 Merge pull request #6568 from aruniiird/upgrade-dependencies-for-operator-sdk-1.x
Bump Controller Runtime version to 0.6
2020-11-19 17:45:18 +01:00
Travis Nielsen 54560250e9 Merge pull request #6497 from sp98/osd-pdb-reconciler
ceph: OSD PDB reconciler changes
2020-11-18 17:30:24 -07:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Sébastien Han a08e08104b Merge pull request #6553 from leseb/mirroring-snap-sched
ceph: add snapshot scheduling for mirrored pools
2020-11-18 15:34:55 +01:00
Santosh Pillai 8602b9c116 ceph: osd pdb reconciler changes
-creates a single PDB (max-unavailable=1) for all OSDs.  This PDB allows one OSD to go down at a given time.
-When a drain is detected, blocking PDBs (max-unavailable=0) will be created for each failure domain that is not being drained and the main PDB (max-unavilable=1) will be deleted. This will allow all the OSDs in the currently drained failure domain to be removed while blocking the deletion  of OSDs in other failure domains.
-Once the PGs are healthy again, the blocking PDBs will be deleted and the main PDB will be restored.
-Add PG healthcheck timeout
-Delete any legacy node drain pods and blocking OSD PDBs

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-11-18 18:25:07 +05:30
Sébastien Han afc7ecff31 ceph: add snapshot scheduling for mirrored pools
Now, we can schedule snapshots on pools from the CephBlockPool CR when
the pool is mirrored.
It can be enabled like this:

```
mirroring:
  enabled: true
  mode: pool
  snapshotSchedules:
    - interval: 24h # daily snapshots
      startTime: 14:00:00-05:00
```

Multiple schedules are supported since snapshotSchedules is a list.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-18 09:50:09 +01:00
subhamkrai 49a3511532 ci: manifests changes to run gh action
to run the integration test there needs
to be a few changes in manifests like using
`deviceFilter` and other related changes.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-11-18 09:31:07 +05:30
subhamkrai a2d3d1d8a4 ceph: update snapshotterVersion to v3.0.0
in the integration test we still use
snapshotteVersion v2.1.0 but it's expected to
use v3.0.0.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-11-16 12:10:50 +05:30
Pete Birley 152a05c85e ceph: update to helm 3 for the rook chart
This updates the chart to make use of helm3 which has been released
for some time, and also permits CRDs to be installed pror to other objects
allowing the chart to be deployed at the same time as CRs for rook objects.

Co-authored-by: Pete Birley <pete@port.direct>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-11 14:41:24 -07:00
Travis Nielsen 61539b6699 Merge pull request #6559 from BlaineEXE/upgrade-test-1.5
ceph: upgrade test Rook only from v1.4 to master
2020-11-11 10:13:39 -07:00
Travis Nielsen 6142efe14b Merge pull request #6579 from subhamkrai/nfs-log
ceph: change nfs-ganesha log level
2020-11-11 10:09:52 -07:00
subhamkrai 7d1b5f7416 ceph: change nfs-ganesha log level
this commit allows changing the default
log level from nfs.yaml file by adding
new command.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-11-11 20:34:07 +05:30
Blaine Gardner 2f034824cb ceph: upgrade test Rook only from v1.4 to master
Since Rook no longer supports legacy filestore devices, there is
no need to keep testing upgrades from v1.2 all the way to master.
We can now just test v1.4 to master.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-11-10 15:44:05 -07:00
Travis Nielsen b9f692a56e ceph: disable the discovery daemon by default
The discovery daemon is not needed in most scenarios, therefore we disable it
by default. More and more clusters are moving to the cluster-on-pvc scenario
which certainly does not need the local discovery. Even where clusters are not
running on PVCs, the discovery is not needed since the device discovery is again
performed in the osd prepare job.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-10 13:16:52 -07:00
Travis Nielsen c029b1513d ceph: remove nullable attribute from pre-1.16 schema
The nullable attribute was causing the integration tests to fail
on 1.13 and earlier where it is not supported in the schema.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit cc8f406ce9)
2020-11-06 10:47:26 -07:00
Sébastien Han e3d032dbc1 Merge pull request #6489 from travisn/stretch-mons
ceph: Stretch cluster configuration
2020-11-06 15:16:39 +01:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
Travis Nielsen b91f4211c9 ceph: configure a stretched cluster
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-05 09:17:52 -07:00
subhamkrai b80d83fcfa ceph: some fields in FS CRD are missing
this commit add the missing `deviceClass` field
in filesystem CRD with replication of pool spec.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-11-05 20:14:11 +05:30
Madhu Rajanna 5de2ded83c ceph: update required rbac for csi
as we are approaching the rook 1.5 release, we
are making a required RBAC changes for cephcsi ahead
of cephcsi release to provide smooth upgrade experience
for the users without any RBAC changes in rook minor release.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-11-04 12:32:08 +05:30
Travis Nielsen d674dcd88c ceph: tests use v14.2.12 due to broken osds in v14.2.13
The ceph v14.2.13 release has breaking changes for the ceph-volume batch
scenario that is preventing non-pvc OSDs from being created. In the short
term we will pin the tests to v14.2.12 and separately we will need to
address the changes needed and/or wait for a fix from ceph.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-03 09:30:07 -07:00
Travis Nielsen 44bf443dca Merge pull request #6495 from LalitMaganti/preserve-fs
ceph: add option to preserve filesystem on CRD deletion
2020-10-30 16:56:38 -06:00
Lalit Maganti c6aec79c4f ceph: add option to preserve filesystem on CRD deletion
Due to #6492, preservePoolsOnDelete is not useful at all for CephFS;
after the filesystem is deleted, the leftover pools cannot be
reassocaited with a newly created filesystem without wiping all
metadata. The only way we can actually preserve data is keeping around
the entire filesystem.

This commit implements a `preserveFilesystemOnDelete` option which work
similar to the existing pool preservation option but instead keeps the
whole CephFS while taking it down and removing all MDSes.

This commit also changes all documentation to refer to this new option
with the intent of essentially deprecating `preservePoolsOnDelete`. IMO,
keeping around `preservePoolsOnDelete` is actively harmful because it
lulls users into thinking their data will be safe but, in reality,
recovering from this situation is highly complex and has large potential
for data loss.

Signed-off-by: Lalit Maganti <lalitm@google.com>
2020-10-30 17:04:18 +00:00
Sébastien Han ea1d71cbfb ceph: add vault kms support for osd encryption
When the Ceph cluster runs on PVC and the OSDs are encrypted we can
store LUKS's Key Encryption Key inside a Key Management System. Today,
Rook only supports HashiCorp Vault: https://www.vaultproject.io/

The CephCluster has now a new "security" field which will plug onto the
KMS. Here is an example:

security:
  kms:
    tokenSecretName: <name of the secret containing a Vault token, used
    to authenticate>
    connectionDetails: < a map of strings containing connection
    information>

Refer to the ceph-cluster-crd documentation to lear more.

Closes: https://github.com/rook/rook/issues/6105
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-30 16:16:33 +01:00
Travis Nielsen 22a1630249 Merge pull request #5483 from cloudandheat/feature/configurable-crush-root
ceph: make root= CRUSH label value configurable
2020-10-29 10:21:02 -06:00
Jonas Schäfer 7ace6ac255 ceph: make root= CRUSH label value configurable
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.

By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.

To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.

The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.

Fixes #4993.

Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>
2020-10-29 09:22:41 +01:00
subhamkraiandtravisn 1b2c15041f ceph: update deprecated CRD apiextensions.k8s.io/v1beta1 to v1
the apiextensions.k8s.io/v1beta1 version of CustomResourceDefinition
is deprecated in Kubernetes v1.16 and will no longer be supported from
v1.19. For now, we changing only for ceph and it's related documented.

Signed-off-by: subhamkrai <srai@redhat.com>
Co-authored-by: travisn <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2020-10-29 06:26:10 +05:30
subhamkrai cb0ca66a6b ci: enable gosec linter in golangci-lint
golangci-lint linter gosec showing more errors
than gosec gh. This commit resolve
new errors. And, removing
nosec comments from autogenerated files.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-10-19 11:50:06 +05:30
tenzen-y 356f21569f ci: cleanup base directory code
Fix the following minor bugs.

- The wrong description of TEST_BASE_DIR environment variable.
- baseTestDir() don't handle error properly

Signed-off-by: tenzen-y <toyonomajyutushi@yahoo.co.jp>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-10-15 14:14:23 +00:00
Takashi IIGUNI 27c7146526 ci: simplify helm path setting
Download the pre-defined helm by default to make the integration test simpler.
As a result, we can make PATH/TEST_HELM_PATH environment variable optional.
In addition, the default helm path can be simplified that doesn't include arch and os.

Signed-off-by: Takashi IIGUNI <iiguni.tks@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-10-15 10:09:00 +09:00