Prometheus found duplicate series for the match group `pool_id`
and `instance` when evaluating the CephPoolGrowthWarning alert
expression. This alert has an evaluation interval of 2 days and
did not consider the changing POD names during the Rook migration process.
This commit adds the `pod` label to the set of considered labels.
Signed-off-by: Matej Feder <matej.feder@dnation.cloud>
The endpoint property honorLabels: true should be set so that
different rgw instances can show separate prometheus
values instead of being aggregated.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Exporter service created by rook operator contains the following labels:
`app=rook-ceph-exporter,rook_cluster=<cluster-name>`
Label `ceph_daemon_id=exporter` is not there and therefore the service
monitor selector can not discover the exporter service.
This fix removes the `ceph_daemon_id=exporter` label from the exporter
service monitor selector.
Signed-off-by: Matej Feder <matej.feder@dnation.cloud>
Pick up the latest ceph prometheus rules from the ceph repo
found at https://github.com/ceph/ceph/blob/master/monitoring/ceph-mixin/prometheus_alerts.yml.
The updates include new rules for monitoring of ceph as well as
adjustments related to the Rook ceph deployment.
List of main adjustments:
- Alerts related to cephadm are excluded
- The PrometheusJobMissing alert is adjusted for the rook-ceph-mgr job, and the PrometheusJobExporterMissing alert is added
Signed-off-by: Matej Feder <matej.feder@dnation.cloud>
including the namespace tag in the monitoring manifests to streamline
the automated process of modifying them through scripting
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
GRPC metrics got deprecated in cephcsi
3.7.0 and the deprecated flags will get
removed in the next release. This PR
removes the deprecated metrics code which
allow us to run with older cephcsi as well.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
To make it easier for the user to switch to a namespace other than the
default 'rook-ceph', the tag '# namespace:cluster' must be inserted in
all the places where this namespace is used. This way, a simple 'sed'
command can update all the Yaml files to the new namespace.
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
The service monitor selectors need to match the mgr service labels.
Since the mgr service labels don't use the mgr_role (only the mgr
service selectors use the mgr_role), the mgr_role should be removed
from the service monitor.
Signed-off-by: travisn <tnielsen@redhat.com>
When the alerts were updated from the ceph repo, the cephadm
alerts were removed, now we also remove two other alerts that
don't apply to rook clusters. The pool growth warning alert
and prometheus job missing alerts need to be removed
as also seen in https://github.com/rook/rook/pull/10109.
Signed-off-by: travisn <tnielsen@redhat.com>
With the new mgr HA implementation (based on readiness probe) we
don't need the mgr sidecar (live-watch) anymore. This commit is
intended to remove all the related code.
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
Predict linear will continue to predict on exporter data from old pods.
We can take the max value for every hour in the 2 day window before we
linearly predict to get the most conservative estimate. This will only
affect the hour in which a new mgr pod is started.
Fixes#10691
Signed-off-by: Jamison Lofthouse <jamison.lofthouse@gmail.com>
Indentation level for the rule array for pools was inconsistent, leading to a broken YAML array.
This commit corrects indentation level for pool rules to align with the rest of the array.
Closes: https://github.com/rook/rook/issues/10653
Signed-off-by: James Harmison <jharmison@gmail.com>
The ceph adm alerts were included in the latest prometheus
rules. With rook these alerts will never be triggered, so
we remove them.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The new alerts picked up for Rook v1.9 directly from the ceph repo
were not all rook-compliant. The alert for the prometheus job not
running is not needed since Rook and K8s go to great lengths to
keep the mgr pod running already.
The alert for the cluster filling up soon is also disabled
until we find how to get the formula working in rook clusters.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The prometheus rules had been previously created if the cephcluster CR
setting monitoring.enabled was set to true. The rules were not customizable
and therefore not flexible enough. Now the rules are installed by the helm
chart. To customize the rules, a post-processor can be applied to the helm
chart.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Even though mention about keda 2.x version, the scaledObject.yaml or
rgw-keda.yaml referring to older keda CRs
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>