Commit Graph
32 Commits
Author SHA1 Message Date
Alexander TrostandTravis Nielsen 36f6807fbe docs: add grafana dashboards files to docs
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2024-09-12 15:44:16 +02:00
Matej Feder 6ba084f35a monitoring: fix CephPoolGrowthWarning expression
Prometheus found duplicate series for the match group `pool_id`
and `instance` when evaluating the CephPoolGrowthWarning alert
expression. This alert has an evaluation interval of 2 days and
did not consider the changing POD names during the Rook migration process.

This commit adds the `pod` label to the set of considered labels.

Signed-off-by: Matej Feder <matej.feder@dnation.cloud>
2024-06-17 09:18:07 +02:00
Travis Nielsen eed1033485 monitoring: set honor labels on the service monitor
The endpoint property honorLabels: true should be set so that
different rgw instances can show separate prometheus
values instead of being aggregated.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2024-06-14 15:38:34 -06:00
Travis Nielsen 02a552ce00 Merge pull request #14312 from matofeder/update-prometheus-alerts
monitoring: update to the latest ceph prometheus rules
2024-06-10 15:17:25 -06:00
Matej Feder 7053e761da monitoring: fix exporter service monitor selector
Exporter service created by rook operator contains the following labels:
`app=rook-ceph-exporter,rook_cluster=<cluster-name>`

Label `ceph_daemon_id=exporter` is not there and therefore the service
monitor selector can not discover the exporter service.

This fix removes the `ceph_daemon_id=exporter` label from the exporter
service monitor selector.

Signed-off-by: Matej Feder <matej.feder@dnation.cloud>
2024-06-07 13:38:15 +02:00
Matej Feder c0a0a3c7b2 monitoring: update to the latest ceph prometheus rules
Pick up the latest ceph prometheus rules from the ceph repo
found at https://github.com/ceph/ceph/blob/master/monitoring/ceph-mixin/prometheus_alerts.yml.

The updates include new rules for monitoring of ceph as well as
adjustments related to the Rook ceph deployment.
List of main adjustments:
- Alerts related to cephadm are excluded
- The PrometheusJobMissing alert is adjusted for the rook-ceph-mgr job, and the PrometheusJobExporterMissing alert is added

Signed-off-by: Matej Feder <matej.feder@dnation.cloud>
2024-06-07 12:47:53 +02:00
Redouane Kachach 4b9e77af1e monitoring: set metrics scraping interval to 10s in examples files
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2024-03-14 16:17:08 +01:00
Redouane Kachach 147f6e2e78 manifest: adding the namespace tag to the monitoring manifests
including the namespace tag in the monitoring manifests to streamline
the automated process of modifying them through scripting

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-11-20 17:19:37 +01:00
Madhu Rajanna 5658e026cb csi: remove deprecated grpc metrics code
GRPC metrics got deprecated in cephcsi
3.7.0 and the deprecated flags will get
removed in the next release. This PR
removes the deprecated metrics code which
allow us to run with older cephcsi as well.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2023-11-07 11:24:17 +01:00
Redouane Kachach ff081cc4ae namespace: adding namespace to all rook-ceph namespaces references
To make it easier for the user to switch to a namespace other than the
default 'rook-ceph', the tag '# namespace:cluster' must be inserted in
all the places where this namespace is used. This way, a simple 'sed'
command can update all the Yaml files to the new namespace.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-10-24 14:17:17 +02:00
travisn 754f073aa5 monitoring: service monitor should not use mgr_role label
The service monitor selectors need to match the mgr service labels.
Since the mgr service labels don't use the mgr_role (only the mgr
service selectors use the mgr_role), the mgr_role should be removed
from the service monitor.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-05-22 14:20:19 -06:00
Redouane Kachach f00bd790a8 mgr: using dynamic mgr_role label to implement mgr HA
Closes: https://github.com/rook/rook/issues/11844

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-03-27 18:38:55 +02:00
Redouane Kachach 456a0c328c Revert "mgr: remove mgr sidecar"
This reverts commit ff75ec96b7.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-03-07 13:12:50 +01:00
travisn fd25eeecf1 monitoring: remove alerts that don't apply to rook
When the alerts were updated from the ceph repo, the cephadm
alerts were removed, now we also remove two other alerts that
don't apply to rook clusters. The pool growth warning alert
and prometheus job missing alerts need to be removed
as also seen in https://github.com/rook/rook/pull/10109.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-03-06 14:46:33 -07:00
Travis Nielsen 2abb8aa08c Merge pull request #11661 from rkachach/fix_issue_removing_sidecar
mgr: remove legacy mgr HA logic and watch-active sidecar
2023-02-21 22:19:30 -07:00
Redouane Kachach ff75ec96b7 mgr: remove mgr sidecar
With the new mgr HA implementation (based on readiness probe) we
don't need the mgr sidecar (live-watch) anymore. This commit is
intended to remove all the related code.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-02-21 17:25:39 +01:00
ahgraber 568623d92b monitoring: remove cephadm ruleset
Remove cephadm ruleset, as they don't apply to rook

Signed-off-by: ahgraber <ahgraber@ninerealmlabs.com>
2023-02-20 21:04:30 -05:00
ahgraber ad2308d655 monitoring: align with ceph/ceph prometheus alerts
Update rook prometheus localrules files to align with ceph/ceph prometheus alerts

Signed-off-by: ahgraber <ahgraber@ninerealmlabs.com>
2023-02-20 13:30:21 -05:00
Avan Thakkar 460900756c core: add service monitor for ceph-exporter service
Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2023-02-15 15:44:56 +05:30
Avan Thakkar 2f8ee60374 core: introduce ceph-exporter
Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2023-02-06 02:47:09 -05:00
parth-gr 703c850b86 core: remove wildcard permission in rbac
Reduce the RBAC scope to the minimum necessary
permissions for rook to operator

Signed-off-by: parth-gr <paarora@redhat.com>
2022-10-10 15:05:56 +05:30
Jamison Lofthouse f553034351 monitoring: fix pool growth warning grouping
Predict linear will continue to predict on exporter data from old pods.
We can take the max value for every hour in the 2 day window before we
linearly predict to get the most conservative estimate. This will only
affect the hour in which a new mgr pod is started.

Fixes #10691

Signed-off-by: Jamison Lofthouse <jamison.lofthouse@gmail.com>
2022-08-11 13:53:28 -07:00
James Harmison 2ae305749a monitoring: fix localrule indent
Indentation level for the rule array for pools was inconsistent, leading to a broken YAML array.
This commit corrects indentation level for pool rules to align with the rest of the array.

Closes: https://github.com/rook/rook/issues/10653
Signed-off-by: James Harmison <jharmison@gmail.com>
2022-07-31 12:31:58 -04:00
Anthony D'Atri dc53497d37 docs: improve descriptions in localrules.yaml
Parallel upstream changes via https://github.com/ceph/ceph/pull/47284

Signed-off-by: Anthony D'Atri <anthonyeleven@users.noreply.github.com>
2022-07-26 16:32:29 -07:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Travis Nielsen 0a25af78dc monitoring: remove cephadm specific alerts
The ceph adm alerts were included in the latest prometheus
rules. With rook these alerts will never be triggered, so
we remove them.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-19 16:03:54 -06:00
Travis Nielsen e45781d0d3 monitoring: disable new alerts that are not applicable to rook
The new alerts picked up for Rook v1.9 directly from the ceph repo
were not all rook-compliant. The alert for the prometheus job not
running is not needed since Rook and K8s go to great lengths to
keep the mgr pod running already.

The alert for the cluster filling up soon is also disabled
until we find how to get the formula working in rook clusters.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-04-19 14:57:05 -06:00
Travis Nielsen 8dd4b77d8a monitoring: update to the latest ceph prometheus rules
Pick up the latest ceph prometheus rules from the ceph repo
found at https://github.com/ceph/ceph/blob/master/monitoring/ceph-mixin/prometheus_alerts.yml.
The updates include many new rules for monitoring of ceph.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-03-21 14:47:56 -06:00
Travis Nielsen 4edcff04f7 monitoring: create prometheus rules with helm chart
The prometheus rules had been previously created if the cephcluster CR
setting monitoring.enabled was set to true. The rules were not customizable
and therefore not flexible enough. Now the rules are installed by the helm
chart. To customize the rules, a post-processor can be applied to the helm
chart.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-03-21 14:47:52 -06:00
Arun Kumar Mohan 054d009af6 monitoring: making alert firing timeout and the doc timeout same
Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2022-02-23 21:15:09 +05:30
Jiffin Tony Thottan e90c36f255 docs: update scaledobject.yaml for applicable with keda v2.5
Even though mention about keda 2.x version, the scaledObject.yaml or
rgw-keda.yaml referring to older keda CRs

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2022-01-18 18:57:28 +05:30
Sébastien Han c890710b63 core: change directory layout
As per discussion, proposing a new layout for the charts/yaml/olm files.

./deploy
├── charts
│   ├── rook-ceph
│   │   └── templates
│   └── rook-ceph-cluster
│       └── templates
├── examples
│   ├── csi
│   │   ├── cephfs
│   │   └── rbd
│   ├── flex
│   ├── monitoring
│   ├── pre-k8s-1.16
└── olm
    └── assemble

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-30 09:12:53 +01:00