Commit Graph
95 Commits
Author SHA1 Message Date
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Yuzuki Mimura 536b59ef0f rgw: change the way to livenessProbe and introduce readinessProbe
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.

Closes: #8407

Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-10-29 15:07:29 +00:00
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
Sébastien Han 8da68bfb78 mon: run ceph commands to mon with timeout
If the mons are not in quorum yet the commands interacting with mon
config store will stale for a very long time.

Closes: https://github.com/rook/rook/issues/8928
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-08 11:57:16 +02:00
Travis Nielsen 0a0b9c98bd build: remove obsolete flex driver
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-09-23 16:17:20 -06:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Travis Nielsen cdfe1982d1 ceph: consolidate the calls to set mon config
The mon config had two different implementations that have evolved
over the lifetime of the project. This is a simple refactor to remove
the SetConfig() option and stick with the MonStore as a single
implementation for updating the mon store.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-08-24 14:38:26 -06:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Yuji Ito d7c46d70d3 ceph: fix overwriting shared livenessProbe
When the Operator set the livenessProbe of the POD, it overwrote the contents of the shared
desiredProbe. Therefore, the wrongly configured livenessProbe was applied to other PODs,
resulting in a CrashLoop. In this PR, I created another instance to avoid overwriting the
desiredProbe.

Closes: https://github.com/rook/rook/issues/8196
Signed-off-by: Yuji Ito <llamerada.jp@gmail.com>
2021-06-28 09:54:33 +00:00
Travis Nielsen b0a63711f5 build: refactor to consolidate the rook.io/v1 package
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-18 19:55:37 -06:00
rohan47 f3d2ae5405 ceph: support specifying only public network or only cluster network
With this change, the Rook ceph cluster can have only a public network or
only a cluster network or both can be specified. Prior to this commit if
only a public network was specified, it was copied to the cluster network.

Signed-off-by: rohan47 <rohgupta@redhat.com>
2021-05-17 17:54:08 +05:30
Sébastien Han 266a359a0e ceph: do not overwrite the probe handler
Changing the livenessprobe handler is not possible anymore for the
simple reason that it does not make sense. Practicly speaking we just
cannot change its value since it is propapaged to all the daemons. The
exec handler contains the socket name of the daemon, because daemons
have different names we cannot possibly make it generic since it will
apply to all.

Closes: https://github.com/rook/rook/issues/7842
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-12 14:06:24 +02:00
Sébastien Han 19de2b28e9 ceph: allow heap dump generation when logCollector is not running
We don't need to set log_file to empty string since "log_to_file" is set
to False. This is probably an old leftover/hack we were doing in the
past when "log_to_file" and other related options were not there.

Because of this line
https://github.com/ceph/ceph/blob/master/src/perfglue/heap_profiler.cc#L99,
Ceph looks for the log_file conf option. In our case it was empty, so
the logging would default to the current directory which points to the
container runtime root when not set and we don't have permission to
write there.

We now force it to Ceph's default so that existing cluster will get the
fix after upgrading.

Just removing the option allows us to get dump in /var/log/ceph.
Phew, what a bug!

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-05-10 17:21:04 +02:00
parth-gr bb23622b72 core: enhancement in debug logs
When DEBUG logging is enabled in the operator, there are a number of overwhelming messages in the log that are very overwhelming and don't seem useful, which makes debug mode difficult to use
Updated and Cutted down the messages that are not useful in debug mode

Closes: https://github.com/rook/rook/issues/7499
Signed-off-by: parth-gr <paarora@redhat.com>
2021-04-28 22:28:10 +05:30
Lars Lehtonen 41c567beaf ceph: fix multiple imports
This fixes additional double-imports.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-13 02:02:09 -07:00
Sébastien Han bce1474f06 ceph: add static IPAM support for CNI
We now support reading Network Attachment Definitions with a "static"
IPAM.
Also added more unit tests.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-25 17:59:34 +01:00
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Sébastien Han a6ef1ff490 ceph: enable pg auto repair
As a recommendation from the core team we should turn on PG auto repair
when running on Bluestore OSDs. The Quincy release will enable this by
default but in the meantime let's enable it here.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-03-12 09:18:00 +01:00
Travis Nielsen 15d0537cec ceph: only raise conditions that represent current state
The conditions on the cephcluster CR were being set to false
when not in progress, which is misleading to the purpose of
conditions. The conditions will now always represent current
state of the cluster. Transient conditions that show progress of
a reconcile will be removed from the status.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-02 14:15:47 -07:00
subhamkrai 4ef4be019b ceph: override default values in liveness probes
this commit set default values or allow overriding of liveness probe
fields if that individual field is 0 or nil.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
2021-02-12 20:35:56 +05:30
ron1 52cf0a3148 ceph: order CephCluster false condition before true Ready condition
The Rancher Continuous Delivery Fleet tool expects a true Ready
condition to be listed after a false condition such as
ProgressingCompleted. This commit orders the false condition
before its corresponding true Ready condition.

Closes: https://github.com/rook/rook/issues/6977
Signed-off-by: ron1 <ron18219@gmail.com>
2021-02-11 12:54:18 -07:00
Sébastien Han c0123cf182 ceph: add cephfs mirroring support
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.

The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050

The same goes for integration tests, they will get added later once we
start testing on Pacific.

Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-28 19:21:18 +01:00
ron1 1e64b82b64 ceph: set ProgressingCompleted CephCluster status condition type
Set ProgressingCompleted CephCluster status condition to include type cephv1.ConditionProgressing

Closes: https://github.com/rook/rook/issues/7055
Signed-off-by: ron1 <ron18219@gmail.com>
2021-01-25 10:09:07 -05:00
ushen d3416fcf2c ceph: allow disable liveness check for mds daemonset
Allow to disable liveness probes for mds pod

Signed-off-by: shenjiatong <yshxxsjt715@gmail.com>
2021-01-11 17:55:51 +08:00
Sébastien Han c6a87203ca ceph: add log collector
We can now collect logs directly into a side-car container.
A new CRD spec has been added:

spec:
  logCollector:
    enabled: true
    periodicity: 24h

Every 24h we will rotate log files for each Ceph daemon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-12-01 16:30:17 +01:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
rohan47 5b5059ec2f ceph: support NAD from diffrent namespaces
Added the support of NAD from diffrent namespaces.
they can be referrenced as <namespace>/<name-of-nad> e.g.,
default/public-nw.
Updated the multus doc to explain the same.

Signed-off-by: rohan47 <rohgupta@redhat.com>
2020-11-17 20:34:58 +05:30
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
subhamkrai 0fddfcf307 ceph: handle golangci-lint linter errcheck error
this commit handle golangci-lint linter errcheck.

`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases

To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-30 22:24:34 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
subhamkrai f9fafe62d4 ceph: handle golangci-lint linter unused
this commit will enable one more linter
in golangci-lint.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 22:53:24 +05:30
Sébastien Han 2dcc6ee392 Merge pull request #5864 from rohan47/whereabouts_fix
ceph: fix networkAttachmentConfig to be compatible with whereabouts cni
2020-08-06 15:11:16 +02:00
Travis Nielsen a3bf935957 ceph: update the cluster status with a clean object
The cluster CR status needs to be updated with the latest object
instead of a cached copy or else it could fail if there is some
other update. Using the Rook clientset ensures it is the latest
instead of using the cached controller runtime object.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-05 15:55:01 -06:00
rohan47 f5bbc8c5c5 ceph: fix networkAttachmentConfig to be compatible with whereabouts cni
In the whereabouts config, the network range is determined by range
option unlike other cni plugins that use Subnet option.

Signed-off-by: rohan47 <rohgupta@redhat.com>
2020-08-05 21:17:13 +05:30
subhamkrai acc4ed5df9 ceph: handling gosec errors that are not checked
this commit handles all the gosec g104
i.e audit errors not checked inside pkg.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-24 20:02:06 +05:30
Sébastien Han 39931b1edf ceph: fix conditions not being applied
Somehow, the client.Update was not updating the cluster object, using
client.Status().Update actually fixed it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-22 17:25:15 +02:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Travis Nielsen 4681f9e73d ceph: consolidate ceph config and client packages
The ceph config and client packages are conceptually the same.
To avoid circular dependencies in some cases, we simplify by combining
the packages into the client package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:43 -06:00
Travis Nielsen 631b13b906 ceph: refactor creds used by operator
The operator should only connect to ceph with a single set of creds.
In a converged cluster this will be the admin creds and in an external
cluster it will be lower-privileged creds. Independent clusters were
implemented with a separate set of creds. To simplify the code these
are now merged to a single set.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:42 -06:00
Sébastien Han ccb52b84e6 ceph: configurable status checks and livenessprobe
This commit allows us to configure status check for each daemon:

* "mon": health check on the ceph monitors (quorum)
* "osd": health check on the ceph osds
* "status": ceph health status check

Each check is controlled by the following settings:

* disabled: whether to disable the check (default: false)
* internal: interval to run the check
* timeout: only valid for mons, is the timeout for unresponsive mon
before failling over.

Example to disable the status health check:

```yaml
healthCheck:
  daemonHealth:
    status:
      disabled: true
```

As part of that, pod's livenessprobe can now be configured via the
following settings:

* disabled: whether to enable or not
* probe: override the current probe in place by a new one

```yaml
healthCheck:
  livenessProbe:
    mon:
      disabled: true
```

Closes: https://github.com/rook/rook/issues/5772
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-10 19:21:59 +02:00
Travis Nielsen 8f9df75f7b ceph: capture logging from the osd daemon
The OSD daemon has been missing critical flags for logging to stderr
where k8s can capture the logs. Without the --log-to-stderr=true,
all the OSD logging was essentially lost until now.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-06 14:38:37 -06:00
Madhu Rajanna 81688398f2 cleanup: use err.Wrap when the formatting is not required
Replaced err.Wrapf with err.Wrap when the formatting
is not required.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-04-29 17:42:59 +05:30
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
n.fraison 6b9a70cb9a ceph: remove specific rgw configuration on mon configuration database
Remove specific rgw config added to the mon configuration database when deleting a CephObjectstore
or when scaling down number of rgw

Signed-off-by: n.fraison <n.fraison@criteo.com>
2020-04-21 22:36:11 +02:00
n.fraison ddadfbc993 ceph: remove leading and trailing space from get config
Value returned by ceph config get can contains some trailing space which leads to
issues when running other ceph command.
For ex.:
2020-04-06 15:32:43.543050 I | exec: Running command: ceph osd pool set test-1.rgw.control pg_num_min 8
 --connect-timeout=15 --cluster=rook-ceph --conf=/var/lib/rook/rook-ceph/rook-ceph.config --keyring=/var/lib/rook/rook-ceph/client.admin.keyring --format json --out-file /tmp/249563089
2020-04-06 15:32:44.005018 I | exec: Error EINVAL: error parsing int value '8

Signed-off-by: n.fraison <n.fraison@criteo.com>
2020-04-06 18:14:50 +02:00
Sébastien Han 20d1543507 ceph: add multus support
You can now use Rook along with Multus. Multus must be up and running
and the right ressources must exist such as NetworkAttachmentDefinition
CR.
The Cluster CR spec has new fields to work with multus:

network:
  provider: multus (or 'host' for hostNetworking)
  selectors:
    public: NetworkAttachmentDefinition name
    cluster: NetworkAttachmentDefinition name

If only a single NetworkAttachmentDefinition is provided Rook will use
both anyway for the Ceph traffic.

Please refer to the doc to learn more.

Closes: https://github.com/rook/rook/issues/4716
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-03 18:00:15 +02:00
Sébastien Han 040193bb5a ceph: remove DaemonType type
This type was a string already and was just making us doing string()
calls all the time to it's not worth it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-01 09:08:18 +02:00
Travis Nielsen e39637930d ceph: set pg_num for rgw metadata pools to lower default
The PG count on metadata pools should default to rgw_rados_pool_pg_num_min
instead of the more general default pg count. This means rgw pools
will default to 8 PGs instead of 32 PGs, which means a lot more pools
can be created before hitting the default PG limit.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-27 09:06:59 -06:00
Travis Nielsen 5254de1a8c ceph: simplify pool model to v1 types
The pools had some legacy structs that translated between the
ceph.v1 types used by the CRDs and the internal implementation
of the pools. This simplifies the pool implementation by removing
the intermediate model and leaving us only with the ceph v1
pool types and no unnecessary translation.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-25 16:59:12 -06:00