This commit uses an csi config map API exposed by Ceph-CSI.
This will reduce the efforts of keeping the configuration in
sync with Ceph-CSI.
Signed-off-by: Praveen M <m.praveen@ibm.com>
Since v18.2.1 is the default recommendation for the ceph version,
update all the docs and examples to that version.
Signed-off-by: travisn <tnielsen@redhat.com>
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
Fixes: #13167
Previously, the mgr did not honor the flag
ContinueUpgradeAfterChecksEvenIfNotHealthy
from the cluster spec. Only osd, mds, and rgw did.
To render the update behavior correct and complete across the daemons, this
change implements the honoring of the flag for the mgr.
Signed-off-by: Michael Adam <obnox@samba.org>
This commit adds new CSIDriverOptions section in
cephCluster CR. This section contains settings
for read affinity and kernel+fuse Mount options
These settings will be injected directly into
rook-ceph-csi-config cm to be applicable per
ceph cluster.
Signed-off-by: Rakshith R <rar@redhat.com>
Restarting the exporter using RollingRelease causes a race condition,
that results in exporter crashing and the ceph health to show a warning.
Signed-off-by: Divyansh Kamboj <dkamboj@redhat.com>
ceph dashboard uses radosgw-admin for certain tasks that
aren't accessible via the rgw REST API. Due to the absence of
a valid ceph.conf file at /etc/ceph/ceph.conf within the mgr pod,
radosgw-admin fails to operate, resulting in 500 errors across
various 'Object Gateway' views on the dashboard. This change
adds CEPH_ARGS environment variable to the mgr pod enabling
its propagation and utilization by the dashboard/radosgw-admin
for executing rgw commands.
closes: https://github.com/rook/rook/issues/13255
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
This patch adds `pgHealthyRegex` field to DisruptionManagementSpec.
`pgHealthyRegex` is a regular expression that is used to determine which
PG states should be considered healthy. The default value of
`pgHealthyRegex` is:
^(active\+clean|active\+clean\+scrubbing|active\+clean\+scrubbing\+deep)$
which is effectively the same as before.
Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
Similar to the ceph crash collector daemon that generates a keyring with more
restrictive privileges, the exporter should also generate and use a more limited keyring.
Signed-off-by: avanthakkar <avanjohn@gmail.com>
During certain maintenance tasks the admin will own running
operations on the ceph mgr, rgw, mds and rbd-mirror daemons
and the operator should not interfere with those operations.
Co-authored-by: gauravsitlani <gaurav.sitlani@live.com>
Signed-off-by: subhamkrai <srai@redhat.com>
when creating networkFence, rbd command was loading
admin config and hence running rbd command use client.admin
in case of external cluster also. With this commit instead
of client.admin user it will use what is being passed to
config.
Signed-off-by: subhamkrai <srai@redhat.com>
This reverts commit fab23d3407.
The mgr requires rw access for the cron job that collects
the crashes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
From the Go specification [1]:
"1. For a nil slice, the number of iterations is 0."
"3. If the map is nil, the number of iterations is 0."
`len` returns 0 if the slice or map is nil [2]. Therefore, checking
`len(v) > 0` before a loop is unnecessary.
[1]: https://go.dev/ref/spec#For_range
[2]: https://pkg.go.dev/builtin#len
Signed-off-by: Eng Zer Jun <engzerjun@gmail.com>
Several messages were showing up in the operator log
more often than necessary, so let's change them to
debug level.
Signed-off-by: travisn <tnielsen@redhat.com>
Using the new configuration parameters prometheusEndpoint and
prometheusEndpointSSLVerify users can configure dashboard to
point to their prometheus instance.
closes: #12876
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
When OSDs flap, ceph stops the OSD daemon if its marked down greater than
5 times in 600 seconds. But OSD pod restarts and marks the OSD `up` again.
This causes the PGs mapped to these OSDs to peer. While the PGs are peering,
IO to these PGs are blocked.
So we need to ensure that if ceph is marking OSD `down` due to flapping, OSD pod
should not restart to mark the OSDs `up` again.
This PR adds a sleep to the OSD pod if the container returned with a 0 exit code
Default behavior is to sleep for 6 hrs. But user can configure it from the
ceph cluster spec.
Signed-off-by: sp98 <sapillai@redhat.com>
While two mgrs is considered sufficient, more mgr daemons are
now allowed in case the admin wants even more fault tolerance
to the active mgr going down, in case multiple mgrs go down.
Up to five mgr pods will be allowed. All mgrs will be in
standby mode except one active mgr. All mgr pods have a sidecar
that will update their respective pod specs with the active
or passive label.
Signed-off-by: travisn <tnielsen@redhat.com>
Previously ceph-exporter would only bind to IPv4 interfaces. Now if the
CephCluster is configured with `dualStack: true` and/or `ipFamily: IPv6`
an additional flag (`--addrs ::`) will be added to the ceph-exporter
container to make it listen both IPv6 and IPv4 interfaces.
Signed-off-by: Matthew Penner <me@matthewp.io>
Change how Rook detects network CIDRs for Multus networks. The IPAM
configuration is only defined as an arbitrary string JSON blob with a
"type" field and nothing more. Rook's detection of CIDRs for whereabouts
had already grown out of date since the initial implementation.
Additionally, Rook did not support DHCP IPAM, which is a reasonable
choice for users. And more, Rook did not support CNI plugin chaining,
which further complicates NADs. Based on the CNI spec, network chaning
can result in any changes to network CIDRs from the first-given plugin.
All these problems make it more and more difficult for Rook to support
Multus by inspecting the NAD itself to predict network CIDRs. Instead,
it is better for Rook to treat the CNI process as a black box. To
preserve legacy functionality of auto-detecting networks and to make
that as robust as possible, change to a canary-style architecture like
that used for Ceph mons, from which Rook will detect the network CIDRs
if possible.
Also allow users to specify overrides for CIDR ranges. This allows Rook
to still support esoteric and unexpected NAD or network configurations
where a CIDR range is not detectable or where the range detected would
be incomplete. Because it may be impossible for Rook to understand the
network CIDRs wholistically while residing only on a portion of the
network, this feature should have been present from Multus's inception.
Improving CIDR auto-detection and allowing users to specify overrides
for auto-detected CIDRs rounds out Rook's Multus support for CephCluster
(core/RADOS) installations. No further architectural changes should be
needed for CephClusters as regards application of public/cluster network
CIDRs for Multus networks.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>