Files
my-rook-config/Documentation/Troubleshooting/performance-profiling.md
T
Blaine Gardner c9d99e01a0 ci: use markdownlint to enforce mkdocs compatibility
mkdocs uses a markdown renderer that is hardcoded to 4 spaces per tab
for detecting indentation levels, including ordered- and
unordered-lists. Since we cannot easily change the renderer, begin using
a markdown linter in CI that will fail if official docs do not adhere to
the spacing rules.

As a starting point, the markdownlint config does not begin with the
default set of checks, which might overwhelm attempts to fix them.
Instead, focus on list-tab-spacing rules and a few other highly useful
checks.

markdownlint also has some gaps in its abilities that allow common Rook
doc issues to pass acceptance. However, it allows creating custom
linting plugins. Create 2 such linting plugins to check 2 things:

- all doc lines (except code blocks) must be aligned to a 4-space
  boundary, without exception. This ensures that markdown will render
  correctly with mkdocs. This unfortunately makes it possible to create
  lists that are internally aligned strangely.
- admonitions must all follow the same format of
  ```
  !!! header
      body
  ```

For the strange lists, this is allowed and renders correctly, but it
looks strange:

```md
- first bullet
- second bullet
    still second bullet
- third bullet

    has a paragraph
    of text inside

- last bullet

Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
2024-04-29 17:25:11 -06:00

2.2 KiB

title
title
Performance Profiling

Collect perf data of a ceph process at runtime

!!! warn This is an advanced topic please be aware of the steps you're performing or reach out to the experts for further guidance.

There are some cases where the debug logs are not sufficient to investigate issues like high CPU utilization of a Ceph process. In that situation, coredump and perf information of a Ceph process is useful to be collected which can be shared with the Ceph team in an issue.

To collect this information, please follow these steps:

  • Edit the rook-ceph-operator deployment and set ROOK_HOSTPATH_REQUIRES_PRIVILEGED to true.
  • Wait for the pods to get reinitialized:
# watch kubectl -n rook-ceph get pods
  • Enter the respective pod of the Ceph process which needs to be investigated. For example:
# kubectl -n rook-ceph exec -it deploy/rook-ceph-mon-a -- bash
  • Install gdb , perf and git inside the pod. For example:
# dnf install gdb git perf -y
  • Capture perf data of the respective Ceph process:
# perf record -e cycles --call-graph dwarf -p <pid of the process>
# perf report > perf_report_<process/thread>
  • Grab the pid of the respective Ceph process to collect its backtrace at multiple time instances, attach gdb to it and share the output gdb.txt:
# gdb -p <pid_of_the_process>

- set pag off
- set log on
- thr a a bt full # This captures the complete backtrace of the process
- backtrace
- Ctrl+C
- backtrace
- Ctrl+C
- backtrace
- Ctrl+C
- backtrace
- set log off
- q (to exit out of gdb)
  • Grab the live coredump of the respective process using gcore:
# gcore <pid_of_the_process>
  • Capture the Wallclock Profiler data for the respective Ceph process and share the output gdbpmp.data generated:
# git clone https://github.com/markhpc/gdbpmp
# cd gdbpmp
# ./gdbpmp.py -p <pid_of_the_process> -n 100 -o gdbpmp.data
  • Collect the perf.data, perf_report, backtrace of the process gdb.txt , core file and profiler data gdbpmp.data and upload it to the tracker issue for troubleshooting purposes.