For nodes that are explicitly requested for deploying
OSDs, they will be skipped temporarily if not
schedulable or not ready. A future reconcile is
expected to schedule them. Allow the OSDs to be
scheduled even on these nodes that are temporarily
down, even if it blocks the reconcile from completing.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
As seen in https://github.com/rook/rook/issues/1988, it's a common
mistake to configure rook nodes with names that don't match Kubernetes's
node label. This PR prints more detailed message to help debugging
problem.
Signed-off-by: Bin Wang <bin.wang@mail.binwang.me>
Add break according to review comment
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Update log message according to review comment
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Adds support for valid arbitrary NodeAffinity JSON input by the user while
maintaining backward compatibility with the older parsing code.
Signed-off-by: Pranshu Srivastava <rexagod@gmail.com>
When a node is being skipped for creation of OSDs, log a message
that indicates whether it was from the node being unschedulable,
the node is not ready, or the node does not meet the placement
criteria specified in the cluster CR.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to k8sutil node functions. By this,
we can handle cancellation during API call of node resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Remove features and design that supports adding OSDs to Ceph clusters
via `spec:driveGroups`. Update the Ceph upgrade doc that informs users
who currently use Drive Groups (we believe there are none of these
users) how to migrate to using the `spec:storage` config.
Resolves https://github.com/rook/rook/issues/7275
Revert "ceph: fix drive group deployment failure"
This reverts commit 76f1d9944e.
Revert "ceph: osd: add drive groups spec to cluster CR"
This reverts commit 7117fc12b7.
Revert "design: ceph orchestrator module add/remove OSDs"
This reverts commit 178187d035.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The osd prepare job detects the failure domain for the OSD by querying
the node topology labels. This topology is then assigned to the OSD
daemon in the CRUSH map. Previously, the affinity was required to be
set in the cluster CR, but it was very difficult to get right. Now the
operator will enforce the correct topology label on the OSD daemon
nodeAffinity by setting the label of the lowest topology in the hierarcy.
For example, if there are region, zone, and rack labels, the rack label
would be used to set the node affinity for the OSD daemon. If an
OSD prepare job is run in rack1, the corresponding OSD daemon will
have node affinity to rack1 to ensure the same topology. Previously,
the OSD could have ended up in rack2 unless the storageClassDeviceSet
placement was very carefully crafted.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
when a node fails some pods remains in terminating while other manage to
be rescheduled on another node successfully.
this commit adds code to clean up all the Rook and CSI pods that are stuck in terminating
state on a failed node.
Signed-off-by: rohan47 <rohgupta@redhat.com>
updating to latest Kubernetes version 1.20.0 fix
security issues. In the current version, it allows
for the token leak in logs when logLevel >= 9.
Signed-off-by: subhamkrai <srai@redhat.com>
Instead of directly defining our own constants, we should be using the k8s
constants for the well-known topology labels.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
this commit handles all the gosec g601
error code (i.e Implicit memory aliasing
of items from a range statement).
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
Add the ability to provision Ceph OSDs with Drive Groups.
This adds Drive Groups to the CephCluster CRD, and it sets code
in place for propagating this config to the OSD provisioning pod.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
The OSDs pick up on several topology labels for CRUSH hierarchy.
The GA label topology.kubernetes.io was partially implemented, but
not picked up by the OSDs. Now the OSDs will pick up both the topology
labels from pre-1.17 such as failure-domain.beta.kubernetes.io/zone
and topology.kubernetes.io/zone.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The rook types used across the storage providers moved from the v1alpha2
package to the v1 package. This commit points the packages at their new
location. Implementation is expected to remain unchanged.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The host name of an OSD in the CRUSH map should be the real
host name for non-portable OSDs. It was incorrectly being set
to the PVC name. Now the non-portable OSDs based on PVCs will
corretly have the host name set to the node name.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Add support for the k8s 1.17 official zone/region failure domain labels.
Our earlier guess that the official labels would be
<failure-domain.kubernetes.io> was wrong; the official is now
<topology.kubernetes.io>.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
The topology of the cluster should be based on the node labels rather
than a setting in the cluster CR. This allows a much richer and more
dynamic topology to be configured. The location will now be ignored
if specified in the cluster CR.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
We removed the topologyAware CRD option since it was redundant with the
use of an OSD being backed by a PVC.
So now, if an OSD is backed by a PVC we assume the topology aware
decision and will discover zone and region labels on that host.
Signed-off-by: Sébastien Han <seb@redhat.com>
The node labels were already supported for zones and regions to add to
the CRUSH map. Now all layers of the CRUSH map will be supported
with the new labels such as topology.rook.io/rack.
The labels will be detected at the osd startup time
similar to the zone and region labels already being detected.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
function AddNodeAffinity was not adding any node
affinity instead it was forming the nodeaffinity
object. renamed it to GenerateNodeAffinity for more
meaningful
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
The node selector for running OSDs on PVCs was using the node name
rather than the node's hostname. In clusters where the node name
is different from the hostname this would cause the osd daemons
to be in pending state indefinitely. Now the node selector will
use the hostname as expected.
Signed-off-by: travisn <tnielsen@redhat.com>
- Added code to support StorageClassDeviceSet spec provided in the cluster-on-pvc.yaml
- The code reads the StorageClassDeviceSet spec and creates pvc based on the ‘count’ field for each device set.
- OSD prepare job is started for each PVC which activates the ceph-volume on each PVC
- Finally OSD is started on each of the PVC device.
Co-authored-by: rohan47 <rohgupta@redhat.com>
Co-authored-by: Ashish Ranjan <aranjan@redhat.com>
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
this patch introduces failure-domain aware scheduling for monitor pods.
previously the only policy was to avoid scheduling monitors on the same
nodes. the new policy tries to spread monitors across zones and nodes.
NOTE: this patch only enforces the well-known k8s "zone"
failure-domain label (i.e. it does not look for region labels). this
decision is made for two reasons. first, eliminating explicit
scheduling is an important goal and will be done as part of upcoming
work. this will enable all supported failure-domain labels to be
handled. second, the "zone" domain is likely more appropriate for a
larger portion of users (as opposed to the geographical distinction
made with region), and supporting multiple failure domain types would
add significant complexity for a temporary resolution to failure
domain scheduling.
at a high-level this patch introduces a data structure `NodeUsage` that
is a pairing of a v1.Node and metadata relevant to scheduling: number of
monitor pods on the node, and its status w.r.t. to being scheduable for
new monitor pods.
each scheduling event computes a unified data structure that organizes
for each node a `NodeUsage` structure into per-failure-zone groups.
subsequent algorithms rely soley on this unified structure to make
scheduling decisions.
there are two entry points to the changes made:
1) during orchestration mon:Cluster:assignMons is invoked to make a
place new monitor pods onto nodes. the core scheduling logic is now
self-contained in mon:scheduleMonitor which implements node and zone
aware scheduling policies.
2) periodic health checks resolve any conflicts that may occur due to
changes to the cluster (e.g. new nodes, configuration changes, etc...).
the two existing checks: mons on invalid nodes, and overloaded nodes are
both handled.
fixes: #2603
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
Previously, Rook Agent and Discovery DaemonSet deployment didn't allow
adding nodeAffinity. This commit adds nodeAffinity spec to daemonSet
deployment, which can be configured through environment variables in
operator deployment yaml.
+ Support multiple LabelKey, each with multiple LabelValue
+ Support multiple LabelKey with no value
Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
This removes the unnecessary check for taints on nodes in the
`GetNodeSchedulable()` function. This fixes that even though the user
has specified `tolerations` for, e.g., `NoSchedule` taints, that would
be ignored and the node(s) directly be "marked" as unschedulable.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
Node validity is not always as simple as a true/false. Refactor each
validity test into its own function for operators that wish to check
validity more granularly. Ceph will use these changes.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>