Prometheus found duplicate series for the match group `pool_id`
and `instance` when evaluating the CephPoolGrowthWarning alert
expression. This alert has an evaluation interval of 2 days and
did not consider the changing POD names during the Rook migration process.
This commit adds the `pod` label to the set of considered labels.
Signed-off-by: Matej Feder <matej.feder@dnation.cloud>
Pick up the latest ceph prometheus rules from the ceph repo
found at https://github.com/ceph/ceph/blob/master/monitoring/ceph-mixin/prometheus_alerts.yml.
The updates include new rules for monitoring of ceph as well as
adjustments related to the Rook ceph deployment.
List of main adjustments:
- Alerts related to cephadm are excluded
- The PrometheusJobMissing alert is adjusted for the rook-ceph-mgr job, and the PrometheusJobExporterMissing alert is added
Signed-off-by: Matej Feder <matej.feder@dnation.cloud>
When the alerts were updated from the ceph repo, the cephadm
alerts were removed, now we also remove two other alerts that
don't apply to rook clusters. The pool growth warning alert
and prometheus job missing alerts need to be removed
as also seen in https://github.com/rook/rook/pull/10109.
Signed-off-by: travisn <tnielsen@redhat.com>
Predict linear will continue to predict on exporter data from old pods.
We can take the max value for every hour in the 2 day window before we
linearly predict to get the most conservative estimate. This will only
affect the hour in which a new mgr pod is started.
Fixes#10691
Signed-off-by: Jamison Lofthouse <jamison.lofthouse@gmail.com>
Indentation level for the rule array for pools was inconsistent, leading to a broken YAML array.
This commit corrects indentation level for pool rules to align with the rest of the array.
Closes: https://github.com/rook/rook/issues/10653
Signed-off-by: James Harmison <jharmison@gmail.com>
The ceph adm alerts were included in the latest prometheus
rules. With rook these alerts will never be triggered, so
we remove them.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The new alerts picked up for Rook v1.9 directly from the ceph repo
were not all rook-compliant. The alert for the prometheus job not
running is not needed since Rook and K8s go to great lengths to
keep the mgr pod running already.
The alert for the cluster filling up soon is also disabled
until we find how to get the formula working in rook clusters.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>