The node usage was used for the obsolete mon selection algorithm, which
was replaced long ago with placement via canary pods.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When the cluster is using HostNetworking, Rook was either picking up the
external or internal IP of the node. This resulted in mon endpoints have
public IP addresses. Those IP are not reachable from within the cluster
so OSD/CSI couldn't access the monitors from the configmap endpoint.
Also, exposing the cluster on a public network does not seem realistic,
so sticky with private/internal IP addresses is better.
Closes: https://github.com/rook/rook/issues/5495
Signed-off-by: Sébastien Han <seb@redhat.com>
this patch adds two new related features. first, it removes the
requirement that all monitor deployments be pinned to nodes using node
selector. this is allowed when not using host networking and when PVC
storage is used. the second feature introduced in this patch is the
elimination of the Rook scheduler for placing monitors that continue to
use the node selector by using canary pods to that are scheduled by k8s,
after which Rook uses the scheduling results to pin the monitors to
nodes.
fixes: #3667
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
By exporting cluster to Cluster we can now reach it from the controller
health check so that it can be populated with 'cluster' data, see the
next commit for better understanding.
Signed-off-by: Sébastien Han <seb@redhat.com>
this patch introduces failure-domain aware scheduling for monitor pods.
previously the only policy was to avoid scheduling monitors on the same
nodes. the new policy tries to spread monitors across zones and nodes.
NOTE: this patch only enforces the well-known k8s "zone"
failure-domain label (i.e. it does not look for region labels). this
decision is made for two reasons. first, eliminating explicit
scheduling is an important goal and will be done as part of upcoming
work. this will enable all supported failure-domain labels to be
handled. second, the "zone" domain is likely more appropriate for a
larger portion of users (as opposed to the geographical distinction
made with region), and supporting multiple failure domain types would
add significant complexity for a temporary resolution to failure
domain scheduling.
at a high-level this patch introduces a data structure `NodeUsage` that
is a pairing of a v1.Node and metadata relevant to scheduling: number of
monitor pods on the node, and its status w.r.t. to being scheduable for
new monitor pods.
each scheduling event computes a unified data structure that organizes
for each node a `NodeUsage` structure into per-failure-zone groups.
subsequent algorithms rely soley on this unified structure to make
scheduling decisions.
there are two entry points to the changes made:
1) during orchestration mon:Cluster:assignMons is invoked to make a
place new monitor pods onto nodes. the core scheduling logic is now
self-contained in mon:scheduleMonitor which implements node and zone
aware scheduling policies.
2) periodic health checks resolve any conflicts that may occur due to
changes to the cluster (e.g. new nodes, configuration changes, etc...).
the two existing checks: mons on invalid nodes, and overloaded nodes are
both handled.
fixes: #2603
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
The desired number of mons could change depending on the number of nodes in a cluster.
For example, three mons would be the min number of mons in a production cluster with at
least three nodes. If there are five or more nodes, the number of mons could increase
to five in order to increase the failure tolerance to two nodes.
This is accomplished by a new setting in the cluster CRD preferredCount. If the number
of hosts exceeds preferredCount, more mons are added to quorum. If the number
of hosts drops below the preferred count, the operator would reduce the quorum size to
the smaller desired count.
Signed-off-by: travisn <tnielsen@redhat.com>
Only the args necessary have traditionally been passed to the mon package.
This has led to many properties in the mon struct, whereas we will
now simply pass the cluster crd spec to simplify.
Signed-off-by: travisn <tnielsen@redhat.com>
The mon operator code is fairly substantial. Separate large files into
more smaller ones with a good separation of names. This will be
important for trying to move all the mon code into the operator.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>