The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Rook checks for down OSDs by checking the `ReadyReplicas` count
in the OSD deployement. When an OSD pod goes into CBLO due to
disk failure, there is a delay before this `ReadyReplicas` count
becomes 0. The deplay is very small but may result in rook missing
OSD down event. As a result no blocking PDBs will be created and
only default PDB with `AllowedDisruptions` count as 0 is available.
This PR tries to solve this. The OSD pdb reconciler will be
reconciled again if `AllowedDisruptions` count in the main
PDB is 0.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit update the PodDisruptionBudget policy to use version v1
Updated to policy/v1 as policy/v1beta1 PodDisruptionBudget is deprecated in v1.21+
Closes: https://github.com/rook/rook/issues/7917
Signed-off-by: parth-gr <paarora@redhat.com>
It's better to create PDB of RGW if ObjectStore's `instances` field is 2.
In the current implementation, at least three desired RGW instances are needed
to create this. If `instances` is 1, we can still safely skip the creation
of PDB because it's useless.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
This PR has following changes around OSD pdbs
- Detect node drains more reliably. When OSD is down and OSD pod is not scheduled to any node or if the scheduled node is not ready, then assume that node is draining.
- Don't set no-out flag on the failure domain if the OSD is down due to reasons other than node drain (say disk failure)
- update unit tests
- update design doc to reflect above changes
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
- Update mon PDB to use maxUnavailable=1 instead of minAvailable. The maxUnavailable will always be 1 irrespective of the number of mons in the cluster
- Move mon PDB reconcile logic from Disruption Controller to Cluster Controller
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Some settinc controller reference code can be replaced
with `controllerutil.SetControllerReference`. It's better to use it
since it provides ownerReference verification.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
This PR makes create and delete events on pdb a no-op. Update event only triggers reconcile if its the main OSD pdb and allowed disruptions is 0.
This prevents controller to reconcile on pdb events namespaces outside rook ceph.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
If the ceph cli outputs an error with "error calling conf_read_file"
this means that the operator has not written its ceph configuration
file. Thus ceph cli commands will fail, so we can just ignore that since
the operator will soon write this file in its initialization sequence.
Signed-off-by: Sébastien Han <seb@redhat.com>
If there is a single mon, we don't expect the PDBs to be created
and neither do we need to log that it is an error condition.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Currently the reconcile of the clusterDisruption controller was triggered mainly due to the
ceph status update in the cephCluster CR. This PR adds reconciles the cluster when:
1. Reconcile when the cluster is created. (This will trigger the first reconcile)
2. Reconcile only when the clusterSpec is updated. (This will avoid triggers when cluster status is updated)
3. Reconcile for events on cephblockpool, cephfilesystem and cephObjectStore.
4. Reconcile for events on Main PDB and when `DisruptionsAllowed` is 0. (that is, when one of the OSD goes down).
5. Reconcile after 30 seconds when there is an active drain going on, that is, pdbStateMap has `draining-failure-domain` as not empty.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
-creates a single PDB (max-unavailable=1) for all OSDs. This PDB allows one OSD to go down at a given time.
-When a drain is detected, blocking PDBs (max-unavailable=0) will be created for each failure domain that is not being drained and the main PDB (max-unavilable=1) will be deleted. This will allow all the OSDs in the currently drained failure domain to be removed while blocking the deletion of OSDs in other failure domains.
-Once the PGs are healthy again, the blocking PDBs will be deleted and the main PDB will be restored.
-Add PG healthcheck timeout
-Delete any legacy node drain pods and blocking OSD PDBs
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
updated mon pdb reconcile to delete mon pdb and create a new one when mon count changes in the cluster spec
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
this commit handle golangci-lint linter staticcheck error.
`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.
To see only `staticcheck` linter output
`golangci-lint run --disable-all -E staticcheck`
Signed-off-by: subhamkrai <srai@redhat.com>
this commit handles all the gosec g601
error code (i.e Implicit memory aliasing
of items from a range statement).
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.
Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
unhandled type assertions can lead to runtime panic. using
the boolean 'ok' flag to handle failure to assert type.
Issue: #4566
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
Previously, delete was returning an error causing the drain-canary
to reconcile even though it no longer existed.
Signed-off-by: Rohan CJ <rohantmp@gmail.com>
The OSDs pick up on several topology labels for CRUSH hierarchy.
The GA label topology.kubernetes.io was partially implemented, but
not picked up by the OSDs. Now the OSDs will pick up both the topology
labels from pre-1.17 such as failure-domain.beta.kubernetes.io/zone
and topology.kubernetes.io/zone.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is to handle all those unhandled errors which raises the gosec warning.
Fixed G104: Unhandled Errors are handled now
Signed-off-by: Nizamudeen <nia@redhat.com>
In the case where the node disappears abrubtly,
there was nothing to trigger the nodedrain reconciler to clean up
the stale canary.
Signed-off-by: Rohan CJ <rohantmp@gmail.com>
Description of your changes:
Fix security issues identified by TrailOfBits
Modifications in osd/status.go:
I opted by remove the offending variables because "c.checkNodesCompleted"
provides this variables always, so no need to initialize them.
var was not used, so i have replaced it by the blank identifier.
Modifications in machinelabel/add.go:
I have followed the same criteria used in other disruption packages,
and return the error if it is produced when we try to "watch" machines.
Closes: https://github.com/rook/rook/issues/4565
[test ceph]
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
When using the "errors" package, using `%+v` (extended format),
each Frame of the error's StackTrace will be printed in detail.
Let's only print `%v` to print the error.
If the error has a Cause it will be printed recursively.
Basically `%+v` has been replaced with `%v` for all `error` type
interfaces, whether the logger is Info, Warning or Error.
Signed-off-by: Sébastien Han <seb@redhat.com>
Earlier, simultaneous drains could allow the drain detection to switch
disabled PDBs before the Ceph health had fully reflected the effects.
Signed-off-by: Rohan CJ <rohantmp@gmail.com>
- Checking whether drain was incorrectly checked precense in map instead
of string length. The actual zero value was an empty string.
- Unsetting drain did not check if the cluster was clean.
Signed-off-by: Rohan CJ <rohantmp@gmail.com>