For nodes that are explicitly requested for deploying
OSDs, they will be skipped temporarily if not
schedulable or not ready. A future reconcile is
expected to schedule them. Allow the OSDs to be
scheduled even on these nodes that are temporarily
down, even if it blocks the reconcile from completing.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Adds support for valid arbitrary NodeAffinity JSON input by the user while
maintaining backward compatibility with the older parsing code.
Signed-off-by: Pranshu Srivastava <rexagod@gmail.com>
When a node is being skipped for creation of OSDs, log a message
that indicates whether it was from the node being unschedulable,
the node is not ready, or the node does not meet the placement
criteria specified in the cluster CR.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to k8sutil node functions. By this,
we can handle cancellation during API call of node resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The osd prepare job detects the failure domain for the OSD by querying
the node topology labels. This topology is then assigned to the OSD
daemon in the CRUSH map. Previously, the affinity was required to be
set in the cluster CR, but it was very difficult to get right. Now the
operator will enforce the correct topology label on the OSD daemon
nodeAffinity by setting the label of the lowest topology in the hierarcy.
For example, if there are region, zone, and rack labels, the rack label
would be used to set the node affinity for the OSD daemon. If an
OSD prepare job is run in rack1, the corresponding OSD daemon will
have node affinity to rack1 to ensure the same topology. Previously,
the OSD could have ended up in rack2 unless the storageClassDeviceSet
placement was very carefully crafted.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
when a node fails some pods remains in terminating while other manage to
be rescheduled on another node successfully.
this commit adds code to clean up all the Rook and CSI pods that are stuck in terminating
state on a failed node.
Signed-off-by: rohan47 <rohgupta@redhat.com>
Instead of directly defining our own constants, we should be using the k8s
constants for the well-known topology labels.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The OSDs pick up on several topology labels for CRUSH hierarchy.
The GA label topology.kubernetes.io was partially implemented, but
not picked up by the OSDs. Now the OSDs will pick up both the topology
labels from pre-1.17 such as failure-domain.beta.kubernetes.io/zone
and topology.kubernetes.io/zone.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The rook types used across the storage providers moved from the v1alpha2
package to the v1 package. This commit points the packages at their new
location. Implementation is expected to remain unchanged.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is to handle all those unhandled errors which raises the gosec warning.
Fixed G104: Unhandled Errors are handled now
Signed-off-by: Nizamudeen <nia@redhat.com>
The minio operator has not had community support nor
any updates since being added to Rook. Support is being
removed from Rook due to this lack of community interest.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The topology of the cluster should be based on the node labels rather
than a setting in the cluster CR. This allows a much richer and more
dynamic topology to be configured. The location will now be ignored
if specified in the cluster CR.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We removed the topologyAware CRD option since it was redundant with the
use of an OSD being backed by a PVC.
So now, if an OSD is backed by a PVC we assume the topology aware
decision and will discover zone and region labels on that host.
Signed-off-by: Sébastien Han <seb@redhat.com>
function AddNodeAffinity was not adding any node
affinity instead it was forming the nodeaffinity
object. renamed it to GenerateNodeAffinity for more
meaningful
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Previously, Rook Agent and Discovery DaemonSet deployment didn't allow
adding nodeAffinity. This commit adds nodeAffinity spec to daemonSet
deployment, which can be configured through environment variables in
operator deployment yaml.
+ Support multiple LabelKey, each with multiple LabelValue
+ Support multiple LabelKey with no value
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
Node validity is not always as simple as a true/false. Refactor each
validity test into its own function for operators that wish to check
validity more granularly. Ceph will use these changes.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>