The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.
Closes: #8407
Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The flex driver has been fully deprecated and thus removed from Rook.
Before upgrading to v1.8, users will need to convert existing flex volumes
from flex to csi volumes.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
The mon config had two different implementations that have evolved
over the lifetime of the project. This is a simple refactor to remove
the SetConfig() option and stick with the MonStore as a single
implementation for updating the mon store.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
When the Operator set the livenessProbe of the POD, it overwrote the contents of the shared
desiredProbe. Therefore, the wrongly configured livenessProbe was applied to other PODs,
resulting in a CrashLoop. In this PR, I created another instance to avoid overwriting the
desiredProbe.
Closes: https://github.com/rook/rook/issues/8196
Signed-off-by: Yuji Ito <llamerada.jp@gmail.com>
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
With this change, the Rook ceph cluster can have only a public network or
only a cluster network or both can be specified. Prior to this commit if
only a public network was specified, it was copied to the cluster network.
Signed-off-by: rohan47 <rohgupta@redhat.com>
Changing the livenessprobe handler is not possible anymore for the
simple reason that it does not make sense. Practicly speaking we just
cannot change its value since it is propapaged to all the daemons. The
exec handler contains the socket name of the daemon, because daemons
have different names we cannot possibly make it generic since it will
apply to all.
Closes: https://github.com/rook/rook/issues/7842
Signed-off-by: Sébastien Han <seb@redhat.com>
We don't need to set log_file to empty string since "log_to_file" is set
to False. This is probably an old leftover/hack we were doing in the
past when "log_to_file" and other related options were not there.
Because of this line
https://github.com/ceph/ceph/blob/master/src/perfglue/heap_profiler.cc#L99,
Ceph looks for the log_file conf option. In our case it was empty, so
the logging would default to the current directory which points to the
container runtime root when not set and we don't have permission to
write there.
We now force it to Ceph's default so that existing cluster will get the
fix after upgrading.
Just removing the option allows us to get dump in /var/log/ceph.
Phew, what a bug!
Signed-off-by: Sébastien Han <seb@redhat.com>
When DEBUG logging is enabled in the operator, there are a number of overwhelming messages in the log that are very overwhelming and don't seem useful, which makes debug mode difficult to use
Updated and Cutted down the messages that are not useful in debug mode
Closes: https://github.com/rook/rook/issues/7499
Signed-off-by: parth-gr <paarora@redhat.com>
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
As a recommendation from the core team we should turn on PG auto repair
when running on Bluestore OSDs. The Quincy release will enable this by
default but in the meantime let's enable it here.
Signed-off-by: Sébastien Han <seb@redhat.com>
The conditions on the cephcluster CR were being set to false
when not in progress, which is misleading to the purpose of
conditions. The conditions will now always represent current
state of the cluster. Transient conditions that show progress of
a reconcile will be removed from the status.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
this commit set default values or allow overriding of liveness probe
fields if that individual field is 0 or nil.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: subhamkrai <srai@redhat.com>
The Rancher Continuous Delivery Fleet tool expects a true Ready
condition to be listed after a false condition such as
ProgressingCompleted. This commit orders the false condition
before its corresponding true Ready condition.
Closes: https://github.com/rook/rook/issues/6977
Signed-off-by: ron1 <ron18219@gmail.com>
With Ceph Pacific, Rook can now deploy the cephfs-mirror daemon.
This initial commit covers the deployment of a single daemon only.
Multiple mirror daemons is currently untested.
Only a single mirror daemon is recommended.
The configuration of peers will come in a later PR since the mgr module
is still pending upstream: https://github.com/ceph/ceph/pull/39050
The same goes for integration tests, they will get added later once we
start testing on Pacific.
Closes: https://github.com/rook/rook/issues/7002
Signed-off-by: Sébastien Han <seb@redhat.com>
We can now collect logs directly into a side-car container.
A new CRD spec has been added:
spec:
logCollector:
enabled: true
periodicity: 24h
Every 24h we will rotate log files for each Ceph daemon.
Signed-off-by: Sébastien Han <seb@redhat.com>
Added the support of NAD from diffrent namespaces.
they can be referrenced as <namespace>/<name-of-nad> e.g.,
default/public-nw.
Updated the multus doc to explain the same.
Signed-off-by: rohan47 <rohgupta@redhat.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
this commit handle golangci-lint linter errcheck.
`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases
To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`
Signed-off-by: subhamkrai <srai@redhat.com>
this commit will enable one more linter ineffassign
in golangci-lint.
This linter throws an error when variable is assigned and never used.
`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
The cluster CR status needs to be updated with the latest object
instead of a cached copy or else it could fail if there is some
other update. Using the Rook clientset ensures it is the latest
instead of using the cached controller runtime object.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In the whereabouts config, the network range is determined by range
option unlike other cni plugins that use Subnet option.
Signed-off-by: rohan47 <rohgupta@redhat.com>
Somehow, the client.Update was not updating the cluster object, using
client.Status().Update actually fixed it.
Signed-off-by: Sébastien Han <seb@redhat.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ceph config and client packages are conceptually the same.
To avoid circular dependencies in some cases, we simplify by combining
the packages into the client package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The operator should only connect to ceph with a single set of creds.
In a converged cluster this will be the admin creds and in an external
cluster it will be lower-privileged creds. Independent clusters were
implemented with a separate set of creds. To simplify the code these
are now merged to a single set.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit allows us to configure status check for each daemon:
* "mon": health check on the ceph monitors (quorum)
* "osd": health check on the ceph osds
* "status": ceph health status check
Each check is controlled by the following settings:
* disabled: whether to disable the check (default: false)
* internal: interval to run the check
* timeout: only valid for mons, is the timeout for unresponsive mon
before failling over.
Example to disable the status health check:
```yaml
healthCheck:
daemonHealth:
status:
disabled: true
```
As part of that, pod's livenessprobe can now be configured via the
following settings:
* disabled: whether to enable or not
* probe: override the current probe in place by a new one
```yaml
healthCheck:
livenessProbe:
mon:
disabled: true
```
Closes: https://github.com/rook/rook/issues/5772
Signed-off-by: Sébastien Han <seb@redhat.com>
The OSD daemon has been missing critical flags for logging to stderr
where k8s can capture the logs. Without the --log-to-stderr=true,
all the OSD logging was essentially lost until now.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.
Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
Remove specific rgw config added to the mon configuration database when deleting a CephObjectstore
or when scaling down number of rgw
Signed-off-by: n.fraison <n.fraison@criteo.com>
Value returned by ceph config get can contains some trailing space which leads to
issues when running other ceph command.
For ex.:
2020-04-06 15:32:43.543050 I | exec: Running command: ceph osd pool set test-1.rgw.control pg_num_min 8
--connect-timeout=15 --cluster=rook-ceph --conf=/var/lib/rook/rook-ceph/rook-ceph.config --keyring=/var/lib/rook/rook-ceph/client.admin.keyring --format json --out-file /tmp/249563089
2020-04-06 15:32:44.005018 I | exec: Error EINVAL: error parsing int value '8
Signed-off-by: n.fraison <n.fraison@criteo.com>
You can now use Rook along with Multus. Multus must be up and running
and the right ressources must exist such as NetworkAttachmentDefinition
CR.
The Cluster CR spec has new fields to work with multus:
network:
provider: multus (or 'host' for hostNetworking)
selectors:
public: NetworkAttachmentDefinition name
cluster: NetworkAttachmentDefinition name
If only a single NetworkAttachmentDefinition is provided Rook will use
both anyway for the Ceph traffic.
Please refer to the doc to learn more.
Closes: https://github.com/rook/rook/issues/4716
Signed-off-by: Sébastien Han <seb@redhat.com>
This type was a string already and was just making us doing string()
calls all the time to it's not worth it.
Signed-off-by: Sébastien Han <seb@redhat.com>
The PG count on metadata pools should default to rgw_rados_pool_pg_num_min
instead of the more general default pg count. This means rgw pools
will default to 8 PGs instead of 32 PGs, which means a lot more pools
can be created before hitting the default PG limit.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The pools had some legacy structs that translated between the
ceph.v1 types used by the CRDs and the internal implementation
of the pools. This simplifies the pool implementation by removing
the intermediate model and leaving us only with the ceph v1
pool types and no unnecessary translation.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>