This commit uses an csi config map API exposed by Ceph-CSI.
This will reduce the efforts of keeping the configuration in
sync with Ceph-CSI.
Signed-off-by: Praveen M <m.praveen@ibm.com>
The mds max daemons had been set arbitrarily to 10.
Allow more mds daemons for clusters that have a
scenario for large numbers of mds to manage the
filesystem metadata.
Signed-off-by: travisn <tnielsen@redhat.com>
This implements the "Ceph Config via Ceph Cluster CRD" design document
as a `cephConfig:` structure on the CRD.
This also fixes the `yq` commands used to manipulate the
`cluster-test.yaml` that caused CI issues for this PR and potentially
unknowingly others.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
this commit adds validating admission policy
for cephcluster cr according to webhook rules.
Not all the webhook can be moved to validating
admission policy for example checking multus
selector validation.
Signed-off-by: subhamkrai <srai@redhat.com>
This commits removes controller-runtime dependencies
from the apis dir and to achieve that we are removing
webhook.
Signed-off-by: subhamkrai <srai@redhat.com>
Originally we create it using this cmd
ceph fs subvolume create <vol_name> <subvol_name>
So we can have 2 variables filesystem and subvolume name,
Currently the CR doesn't allow us to make subvolume-name
as constant as needed to "csi" because of k8s limitations
Signed-off-by: parth-gr <paarora@redhat.com>
This commit adds new CSIDriverOptions section in
cephCluster CR. This section contains settings
for read affinity and kernel+fuse Mount options
These settings will be injected directly into
rook-ceph-csi-config cm to be applicable per
ceph cluster.
Signed-off-by: Rakshith R <rar@redhat.com>
This patch adds `pgHealthyRegex` field to DisruptionManagementSpec.
`pgHealthyRegex` is a regular expression that is used to determine which
PG states should be considered healthy. The default value of
`pgHealthyRegex` is:
^(active\+clean|active\+clean\+scrubbing|active\+clean\+scrubbing\+deep)$
which is effectively the same as before.
Signed-off-by: Ryotaro Banno <ryotaro.banno@gmail.com>
Allow end-user to have full control on NFS-Ganesha liveness probe:
- Disabled
- Enabled, using user-provided definition
- Enabled, using Rook's default liveness probe for Ganesha.
Signed-off-by: Shachar Sharon <ssharon@redhat.com>
The merge of the go modules in the apis subdirectory
missed some changes from another merge around the same
time, so this will add the missing updates and
fix the CI.
Signed-off-by: travisn <tnielsen@redhat.com>
To keep the go dependencies minimal, add the go packages
just for the pkg/apis directory. Projects that only reference
the rook apis can then have a smaller footprint for Go
dependencies.
Signed-off-by: travisn <tnielsen@redhat.com>
Currently, the function receives a single namespace string and appends
it to the users. This approach is insufficient when dealing with
multiple CephClusters running in different namespaces, as there will be
only one SCC. This update modifies the function to create users for all
specified namespaces, ensuring compatibility with multiple namespaces.
Signed-off-by: Nitin Goyal <nigoyal@redhat.com>
Using the new configuration parameters prometheusEndpoint and
prometheusEndpointSSLVerify users can configure dashboard to
point to their prometheus instance.
closes: #12876
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
When OSDs flap, ceph stops the OSD daemon if its marked down greater than
5 times in 600 seconds. But OSD pod restarts and marks the OSD `up` again.
This causes the PGs mapped to these OSDs to peer. While the PGs are peering,
IO to these PGs are blocked.
So we need to ensure that if ceph is marking OSD `down` due to flapping, OSD pod
should not restart to mark the OSDs `up` again.
This PR adds a sleep to the OSD pod if the container returned with a 0 exit code
Default behavior is to sleep for 6 hrs. But user can configure it from the
ceph cluster spec.
Signed-off-by: sp98 <sapillai@redhat.com>
While two mgrs is considered sufficient, more mgr daemons are
now allowed in case the admin wants even more fault tolerance
to the active mgr going down, in case multiple mgrs go down.
Up to five mgr pods will be allowed. All mgrs will be in
standby mode except one active mgr. All mgr pods have a sidecar
that will update their respective pod specs with the active
or passive label.
Signed-off-by: travisn <tnielsen@redhat.com>
Change how Rook detects network CIDRs for Multus networks. The IPAM
configuration is only defined as an arbitrary string JSON blob with a
"type" field and nothing more. Rook's detection of CIDRs for whereabouts
had already grown out of date since the initial implementation.
Additionally, Rook did not support DHCP IPAM, which is a reasonable
choice for users. And more, Rook did not support CNI plugin chaining,
which further complicates NADs. Based on the CNI spec, network chaning
can result in any changes to network CIDRs from the first-given plugin.
All these problems make it more and more difficult for Rook to support
Multus by inspecting the NAD itself to predict network CIDRs. Instead,
it is better for Rook to treat the CNI process as a black box. To
preserve legacy functionality of auto-detecting networks and to make
that as robust as possible, change to a canary-style architecture like
that used for Ceph mons, from which Rook will detect the network CIDRs
if possible.
Also allow users to specify overrides for CIDR ranges. This allows Rook
to still support esoteric and unexpected NAD or network configurations
where a CIDR range is not detectable or where the range detected would
be incomplete. Because it may be impossible for Rook to understand the
network CIDRs wholistically while residing only on a portion of the
network, this feature should have been present from Multus's inception.
Improving CIDR auto-detection and allowing users to specify overrides
for auto-detected CIDRs rounds out Rook's Multus support for CephCluster
(core/RADOS) installations. No further architectural changes should be
needed for CephClusters as regards application of public/cluster network
CIDRs for Multus networks.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The object user was previously required to be created in the
same namespace as the object store and the cluster. Now,
the object user can be reconciled even in a different namespace
from the cluster and object store. The namespace would be specified
in the object user CR.
Signed-off-by: travisn <tnielsen@redhat.com>
The go package dependencies (go.mod and go.sum) were created
under the pkg/apis directory to simplify the dependencies for
other projects referencing the rook repo. The downside is that
the dependabot can no longer open a working and valid PR since
the bot is not aware of the go.mod and go.sum in the apis directory.
Since the reduction of dependencies for vault in #12455,
the list of extra dependencies is not quite so extensive.
We are putting effort into shrinking the dependency list instead
of using the modules files in the apis subdirectory.
Note that #12419 updated the dependencies for the latest controller
runtime which also increased the size of the modules
list in the pkgs subdirectory, which means the difference
in dependencies in the pkgs subdirectory is no longer
signficant anyway.
Signed-off-by: travisn <tnielsen@redhat.com>
If osd store is updated in the ceph cluster, then
delete OSDs one by one, cleanup disks and provision a new OSD on
the same disk
Signed-off-by: sp98 <sapillai@redhat.com>