v1 version of admission controller require minimum v1.16.0
of k8s. So, this commits disable admission controller when
k8s version older than v1.16.0.
Signed-off-by: subhamkrai <srai@redhat.com>
disabling admission controller for the upgrade test.
Upgrade test was using older controller runtime version
and api v1 need latest version controller runtime(v0.7).
Signed-off-by: subhamkrai <srai@redhat.com>
The cockroachDB operator has not had community support in Rook.
Therefore, the time has come to deprecate and remove it.
If the sources are still needed, there is always git history.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The integration tests check after every test suite that the cluster is
either ready or connected. Sometimes the operator is in progressing state
because the reconcile may be triggered by some event that is unpredictable
at the end of the test. For test stability we allow the progressing status.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With the 1.5 schema added to the CRDs, the devices at the root
level of the storage element were missed. Now the ability to
specify devices at the root storage level to apply to all
nodes is restored.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When verifying OBC creation, validating that secret and configmap exist
is good, but the definitive validation is to check that the OBC's phase
is "Bound".
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Update to the latest lib bucket provisioner code.
Fixes issue 6650
Modifies CRD for objectbucketclaims to fix an additional bug where an
ObjectBucket's 'ClaimRef' is lost due to the CRD validation being
specified incorrectly.
Changes OBC deletion/cleanup to delete the bucket before the user. A
user cannot be deleted without an unsafe purge option if the user has
buckets associated to it.
Does not reintroduce bug 6767 from previous fix for 6650
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The latest ceph-volume does not support creating OSDs on partitions.
The github actions only have a partition available, so we will run
the github tests on v14.2.12 and v15.2.7 that still support partitions,
while the Jenkins environment has a raw device available where we
can run the latest versions of Ceph in the tests.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The snapshot controller is specifying to use the canary image instead
of the expected version. Update the tests to set the correct version.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The hostpath provisioner has never been supported and neither is it
working on K8s 1.20. For the tests we will simply use static
local PVs so we can avoid the whole problem of an unsupported
hostpath provisioner.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Add meta-comments to manifests to allow basic templating via `sed`.
Add the following types of meta-comments:
- # namespace:X
- A basic namespace
- replace the field with a namespace for X
- # serviceaccount:namespace:X
- A service account namespace for SCC
- e.g., "system:serviceaccount:<ns>:rook-ceph-system"
- Replace the namespace "<ns>" with with a namespace for X
- # provisioner:namespace:X
- A provisioner identifier with namespace prefix
- e.g., "<ns>.cephfs.csi.ceph.com"
- e.g., "<ns>.ceph.rook.io/bucket"
- Replace the namespace "<ns>" with a namespace for X
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
When the cluster is external we want to expose manager service port
along with the endpoints.
This allows us to connect but Rook will keep on using 9283 as a facing
port for Prometheus and more.
```
monitoring:
enabled: true
...
...
externalMgrPrometheusPort: 9283
```
Signed-off-by: Sébastien Han <seb@redhat.com>
testCephHelmSuite is failing in pre-k8s v1.16
due to invalid schema. v1beta1 doesn't support
null value so returning quotes instead of an empty
string.
Signed-off-by: subhamkrai <srai@redhat.com>
We can now collect logs directly into a side-car container.
A new CRD spec has been added:
spec:
logCollector:
enabled: true
periodicity: 24h
Every 24h we will rotate log files for each Ceph daemon.
Signed-off-by: Sébastien Han <seb@redhat.com>
Rook's crashcollector pod posts entries to the ceph cluster when a crash occurs.
Over time the number cluster may hold crash entries needlessly.
To clean up old crash entries, this PR adds a field to the ceph cluster CR for the user to specify the number of days a crash entry should be kept for.
Providing a value for the field keepXDays creates a cronjob that runs every day at midnight, calling "ceph crash prune <keepXDays>".
Closes: https://github.com/rook/rook/issues/6332
Signed-off-by: Renan Campos <rcampos@redhat.com>
-creates a single PDB (max-unavailable=1) for all OSDs. This PDB allows one OSD to go down at a given time.
-When a drain is detected, blocking PDBs (max-unavailable=0) will be created for each failure domain that is not being drained and the main PDB (max-unavilable=1) will be deleted. This will allow all the OSDs in the currently drained failure domain to be removed while blocking the deletion of OSDs in other failure domains.
-Once the PGs are healthy again, the blocking PDBs will be deleted and the main PDB will be restored.
-Add PG healthcheck timeout
-Delete any legacy node drain pods and blocking OSD PDBs
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Now, we can schedule snapshots on pools from the CephBlockPool CR when
the pool is mirrored.
It can be enabled like this:
```
mirroring:
enabled: true
mode: pool
snapshotSchedules:
- interval: 24h # daily snapshots
startTime: 14:00:00-05:00
```
Multiple schedules are supported since snapshotSchedules is a list.
Signed-off-by: Sébastien Han <seb@redhat.com>
to run the integration test there needs
to be a few changes in manifests like using
`deviceFilter` and other related changes.
Signed-off-by: subhamkrai <srai@redhat.com>
This updates the chart to make use of helm3 which has been released
for some time, and also permits CRDs to be installed pror to other objects
allowing the chart to be deployed at the same time as CRs for rook objects.
Co-authored-by: Pete Birley <pete@port.direct>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Since Rook no longer supports legacy filestore devices, there is
no need to keep testing upgrades from v1.2 all the way to master.
We can now just test v1.4 to master.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The discovery daemon is not needed in most scenarios, therefore we disable it
by default. More and more clusters are moving to the cluster-on-pvc scenario
which certainly does not need the local discovery. Even where clusters are not
running on PVCs, the discovery is not needed since the device discovery is again
performed in the osd prepare job.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The nullable attribute was causing the integration tests to fail
on 1.13 and earlier where it is not supported in the schema.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
(cherry picked from commit cc8f406ce9)
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
as we are approaching the rook 1.5 release, we
are making a required RBAC changes for cephcsi ahead
of cephcsi release to provide smooth upgrade experience
for the users without any RBAC changes in rook minor release.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
The ceph v14.2.13 release has breaking changes for the ceph-volume batch
scenario that is preventing non-pvc OSDs from being created. In the short
term we will pin the tests to v14.2.12 and separately we will need to
address the changes needed and/or wait for a fix from ceph.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Due to #6492, preservePoolsOnDelete is not useful at all for CephFS;
after the filesystem is deleted, the leftover pools cannot be
reassocaited with a newly created filesystem without wiping all
metadata. The only way we can actually preserve data is keeping around
the entire filesystem.
This commit implements a `preserveFilesystemOnDelete` option which work
similar to the existing pool preservation option but instead keeps the
whole CephFS while taking it down and removing all MDSes.
This commit also changes all documentation to refer to this new option
with the intent of essentially deprecating `preservePoolsOnDelete`. IMO,
keeping around `preservePoolsOnDelete` is actively harmful because it
lulls users into thinking their data will be safe but, in reality,
recovering from this situation is highly complex and has large potential
for data loss.
Signed-off-by: Lalit Maganti <lalitm@google.com>
When the Ceph cluster runs on PVC and the OSDs are encrypted we can
store LUKS's Key Encryption Key inside a Key Management System. Today,
Rook only supports HashiCorp Vault: https://www.vaultproject.io/
The CephCluster has now a new "security" field which will plug onto the
KMS. Here is an example:
security:
kms:
tokenSecretName: <name of the secret containing a Vault token, used
to authenticate>
connectionDetails: < a map of strings containing connection
information>
Refer to the ceph-cluster-crd documentation to lear more.
Closes: https://github.com/rook/rook/issues/6105
Signed-off-by: Sébastien Han <seb@redhat.com>
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.
By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.
To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.
The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.
Fixes#4993.
Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>