Commit Graph
33 Commits
Author SHA1 Message Date
Deepika Upadhyay c5d27467d6 object: add rgw ops sidecar for op logs
the rgw operations for s3 can now be accessible using sidecar
rgw-ops-log availabe in json form that can be further filtered logging
for observability, this will set the rgw_enable_ops_log setting

Signed-off-by: Deepika Upadhyay <deepika.upadhyay@clyso.com>
2024-12-17 22:36:03 +05:30
df511fb58f ci: update golangci-lint to the latest version (v1.62)
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.

Additionally, it  silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
 and string format errors found by golangci-lint, while at it.

Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
2024-12-14 14:47:30 +01:00
travisn 03d077aa6b core: remove support for ceph pacific
Pacific is end of life and no longer necessary to
support in Rook with v1.13.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-11-14 17:07:03 -07:00
sp98 4eb9f62205 osd: make osd pod to sleep when osds are flapping
When OSDs flap, ceph stops the OSD daemon if its marked down greater than
5 times in 600 seconds. But OSD pod restarts and marks the OSD `up` again.
This causes the PGs mapped to these OSDs to peer. While the PGs are peering,
IO to these PGs are blocked.

So we need to ensure that if ceph is marking OSD `down` due to flapping, OSD pod
should not restart to mark the OSDs `up` again.

This PR adds a sleep to the OSD pod if the container returned with a 0 exit code
Default behavior is to sleep for 6 hrs. But user can configure it from the
ceph cluster spec.

Signed-off-by: sp98 <sapillai@redhat.com>
2023-09-15 22:59:46 +05:30
subhamkrai 29d2b6a071 core: restart ceph daemons when network updated
We need to restart all the ceph daemons whenever
cephCluster network settings are modified like
requiremsgr2, encryption and compression. This
required for Ceph to consider the new settings
it require new ceph daemons all over.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-09-01 11:30:48 +05:30
avanthakkar 541d091f9c core: use ROOK_CEPH_MON_HOST from config store in OSD pods too
Signed-off-by: avanthakkar <avanjohn@gmail.com>

Volume "ceph-daemons-sock-dir" is coming empty in case if dataDirHostPath,
which is the case for osd onPVC. Fix the volume creation by using the
ceph cluster spec dataDirHostPath, which allows to run socket commands
on osd containers.
2023-06-01 21:09:58 +05:30
Javier 02e17196f2 core: use -default-* flags
enable flags with --default prefix for --log-to-stderr, --mon-cluster-log-to-stderr, --err-to-stderr, and --log-stderr-prefix

Signed-off-by: Javier <sjavierlopez@gmail.com>
2023-05-30 11:52:26 -06:00
Redouane Kachach 0ae7867dd1 Revert "mgr: use k8s readiness probe to implement mgr HA"
This reverts commit dc76f81fea.

Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-03-07 13:12:50 +01:00
Travis Nielsen b008c5753f Merge pull request #11643 from rkachach/fix_issue_11640
mgr: use k8s readiness probe to implement mgr HA
2023-02-15 09:51:48 -07:00
Redouane Kachach dc76f81fea mgr: use k8s readiness probe to implement mgr HA
The idea behind this change is to use the readiness probe to implement
the mgr HA mechanism. In the current ceph mgr implementation only the
active instance offers the command 'mgr_status' through the admin
socket. We use this command combined with a Readiness Exec Probe to
detect which manager is active. Kubernetes will automatically mark
it as ready and redirect any service traffic to the active instance.

Closes: https://github.com/rook/rook/issues/11640
Closes: https://github.com/rook/rook/issues/11638
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
2023-02-15 10:12:16 +01:00
subhamkrai 036715c3b2 rbdmirror: set log rotation from 7 to 4x i.e 28
increasing the rotation from default 7 to 28 as
in case of rbdmirror logs seems not enough in some cases
with maxLogSize 500 so it's better to increase the rotation
for rbdmirror specific.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-02-13 21:37:57 +05:30
Avan Thakkar 934aa91056 operator: make imagePullPolicy customizable for csi driver and ceph pods
Introduce a new env variable ROOK_CSI_IMAGE_PULL_POLICY in rook operator configmap which should be used to
customize the imagePullPolicy for the csi driver and imagePullPolicy property in cephVersionSpec for ceph pods.

Signed-off-by: Avan Thakkar <athakkar@redhat.com>
2022-09-27 11:57:43 +05:30
motorailgun ff0951738f operator: improve ProbeHandler error message
This commit implements more diagnostic error for unsccessful
Liveness- and Readiness-Probe of Pods.

Added codes are expected to catch failures of
`ceph status` and `ceph mon_status`, and report it.

Closes: https://github.com/rook/rook/issues/9846
Closes: https://github.com/rook/rook/issues/9852

Signed-off-by: motorailgun <motoi_public@mail.aria-on-the-planet.es>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2022-09-07 07:35:33 +00:00
subhamkrai 62f73dcd98 core: add support to rotate log based on logfile size
this commits add new field `MaxLogSize` inside `LogCollectorSpec` of
cephCluster cr which will take max size of log after which we want to
rotate the log.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-07-26 13:06:39 +00:00
Travis Nielsen dad97f3425 core: remove support for ceph octopus
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-07-07 15:03:26 -06:00
subhamkrai 62f0fb2b1d core: increase liveness probe timeout to 2s
we have noticed multiple failures because of the probe
failing, most of the time it's due to fewer resources.
But increasing timeout fixes that, so increasing the
probe timeout to 2s from default 1s so that it will
give more time to probe before failing.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-06-17 19:06:40 +05:30
Alexander Trost 8686296e17 core: remove double imported packages
This removes double package imports. Example:
```
"github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
cephv1 "github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
```
Only one is now being used as shown in go-staticcheck ST1019

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-25 13:51:45 +02:00
subhamkrai bf7daccf60 core: make code changes to support latest cntrl runtime
making necessary code changes to support controller
runtime version.

Signed-off-by: subhamkrai <srai@redhat.com>
2022-04-20 22:30:15 +05:30
Travis Nielsen 4edcff04f7 monitoring: create prometheus rules with helm chart
The prometheus rules had been previously created if the cephcluster CR
setting monitoring.enabled was set to true. The rules were not customizable
and therefore not flexible enough. Now the rules are installed by the helm
chart. To customize the rules, a post-processor can be applied to the helm
chart.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-03-21 14:47:52 -06:00
Blaine Gardner c07d89d9ea core: rgw: allow specifying daemon startup probes
Allow specifying daemon startup probes where we also allow configuring
liveness probes. Startup probes allow Rook to tolerate when Ceph daemons
occasionally take a long time to start up while not also making
Kubernetes liveness probes slower to detect runtime failures of daemons.

Startup probes are beta in Kubernetes 1.18, so we should not enable
probes by default for earlier Kubernetes versions.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-12-21 15:12:41 -07:00
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Sébastien Han b578f916e7 ceph: add fs mirror config
Similarly to block volume replication, Ceph is capable of replicating the
content of a Ceph Filesystem from one cluster to another.
For this, during the 1.6 cycle, we introduced a new CRD called
CephFilesystemMirror which effectively deploys a cephfs-mirror daemon.
However, configuring peers to enable replication between two clusters
had to be done manually.
Also various bug fix made it in Ceph eventually and the minimum required
version for this to work is to run on Ceph Pacific 16.2.5 at least.

So the automatic configuration of Ceph Filesystem peers is now possible.

By editing the CephFilesystem CRD, you can now turn on mirroring:

```yaml
  mirroring:
    enabled: false
    # list of Kubernetes Secrets containing the peer token
    # for more details see: https://docs.ceph.com/en/latest/dev/cephfs-mirroring/#bootstrap-peers
    peers:
      secretNames:
        - secondary-cluster-peer
```

Also, the mirroring status is displayed in the CR status:

```
status:
  info:
    fsMirrorBootstrapPeerSecretName: fs-peer-token-myfs
  mirroringStatus:
    daemonsStatus:
    - daemon_id: 4186
      filesystems:
      - filesystem_id: 2
        name: myfs
    lastChecked: "2021-07-01T14:16:29Z"
  phase: Ready
  snapshotScheduleStatus:
    lastChecked: "2021-07-01T14:16:29Z"
    snapshotSchedules:
    - fs: myfs
      path: /
      rel_path: /
      retention: {}
      schedule: 24h
```

Closes: https://github.com/rook/rook/issues/7063
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 17:35:19 +02:00
Sébastien Han 5e1e9f4c92 ceph: actively update the service endpoint for external mgr
If the cluster is external we want to periodically rehydrate the mgr
endpoint. This handles the scenarion where the active manager changes,
 so we need to update the endpoint with the new IP address.
The create-external-cluster-resources.py script now requires an extra
permission to query the manager service so Rook can discover the active
one and its IP.

Testing:

```
[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
    "epoch": 37,
    "available": true,
    "active_name": "b",
    "num_standby": 1
}

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME                     ENDPOINTS          AGE
rook-ceph-mgr-external   172.17.0.12:9283   3m10s

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] k scale --replicas=0 deployment rook-ceph-mgr-b
deployment.apps/rook-ceph-mgr-b scaled

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] minikube kubectl -- exec -n rook-ceph deploy/rook-ceph-tools -ti -- ceph mgr stat
{
    "epoch": 40,
    "available": true,
    "active_name": "a",
    "num_standby": 0
}

[leseb@tarox~/go/src/github.com/rook/rook][external-active-mgr-change] kubectl -n rook-ceph-external get ep
NAME                     ENDPOINTS          AGE
rook-ceph-mgr-external   172.17.0.13:9283   3m55s
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-11 16:08:40 +02:00
Lars Lehtonen 41c567beaf ceph: fix multiple imports
This fixes additional double-imports.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-13 02:02:09 -07:00
Sébastien Han 51e7310e09 ceph: add dualstack support on pacific
With Pacific comes the support for dualstack where ceph daemons can
listen on both ipv4 and ipv6 stacks.
A new field in the network spec has been added: `dualStack`

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-31 18:04:29 +02:00
Santosh Pillai 70c5f70701 ceph: support IPv6 single-stack for ceph
Add --ms-bind-ipv6 true args to mom, osd, mgr, object, mds, rbd daemons.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2020-10-02 01:03:03 +05:30
Travis Nielsen ce92249725 ceph: continue with memory limits below min settings
The operator will now allow the resource limits to be applied below the
recommended minimums. In small clusters, even the recommended minimums
may not be necessary. A warning is still printed to the operator log,
but we allow the configuration to continue.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-19 11:07:16 -06:00
Sébastien Han c8de7009de ceph: increase liveness probe start delay on OSD
When the cluster is loaded and we restart an OSD, it will need some time
to respond to socket calls, basically more to be ready.
Increasing the initialDelaySeconds of the liveness probe fixes that
issue. For OSD, it waits for 45 sec, where other daemons 10 sec.

Closes: https://github.com/rook/rook/issues/5492
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-05 15:53:49 +02:00
Sébastien Han 040193bb5a ceph: remove DaemonType type
This type was a string already and was just making us doing string()
calls all the time to it's not worth it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-01 09:08:18 +02:00
Sébastien Han 7788901a88 ceph: add liveness probe to mon, mds and osd daemons
Now Kubernetes will perform liveness checks on mon, mds and osd daemons.
The command will:

* call the socket (check for existence)
* execute a command and check the return code (success if 0)

This handles the case where the daemon is stuck locally and
unresponsive. It's unlikely but not impossible.
These checks bring more robustness to the implementation.

rbd-mirror and nfs have been leftover for the following reason. The
rbd-mirror socket name is different from other daemons (could be fixed
though): /run/ceph/ceph-client.rbd-mirror.a.1.94362516231272.asok also,
the command to call would need to be changed from "status" to "rbd
mirror status" so we can keep this for a later.
The nfs ganesha has no socket only a PID file which doesn't mean much.
No PID means the process does not run so Kubernetes will already handle
this and the pod will crash loop.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-31 15:20:42 +02:00
Sébastien Han a3068dee0b ceph: separate controller for CephBlockPool CRD
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-02 17:31:03 +01:00