few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
This fixes the the MultiClusterDeploySuite CI failure. The operator
is timing out waiting for the mgr deployments to be
ready, according to the WaitForDeploymentToStart() method. After
the reconcile times out after about five minutes, the next
reconcile succeeds since the wait is only done for new mgr
deployments.
Closes: https://github.com/rook/rook/issues/11685
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
The idea behind this change is to use the readiness probe to implement
the mgr HA mechanism. In the current ceph mgr implementation only the
active instance offers the command 'mgr_status' through the admin
socket. We use this command combined with a Readiness Exec Probe to
detect which manager is active. Kubernetes will automatically mark
it as ready and redirect any service traffic to the active instance.
Closes: https://github.com/rook/rook/issues/11640
Closes: https://github.com/rook/rook/issues/11638
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
increasing the rotation from default 7 to 28 as
in case of rbdmirror logs seems not enough in some cases
with maxLogSize 500 so it's better to increase the rotation
for rbdmirror specific.
Signed-off-by: subhamkrai <srai@redhat.com>
The crash collector controller is designed for watching nodes where
ceph daemons are running, and ensuring a special daemon is running
on that node to provide additional support for ceph on that node.
The crash collector is the first example of a daemon that should be
running on all the ceph daemon nodes. The next example of such a
node daemon will be the ceph exporter that will listen for the
ceph metrics as described in the design doc.
https://github.com/rook/rook/blob/master/design/ceph/ceph-exporter.md
Now the crash collector controller is renamed to the node daemon controller
so the ceph exporter daemon can also be managed by the same controller.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.
This feature is intended to be used for testing purposes only.
Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
The mons that are out of quorum may cause ceph commands to timeout
or fail unnecessarily trying to connect to a mon that is no longer
online. Now the mon health check will update the mon endpoints configmap
when a mon is detected out of quorum. This also means that if the
operator is restarted during a mon failover, the failed mon will no
longer remain in the ceph.conf, thus allowing the quorum to be more
likely to respond to the mons that are still in quorum.
The update for mons out of quorum only applies if other mons are in
quorum. If quorum is down, the configmap will not keep track of the offline
mons since too many are offline.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The crash collector does not have the command line arguments
to run as ceph user id 167, so we set the security context
to run as the ceph user in the main crash collector
container.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Introduce a new env variable ROOK_CSI_IMAGE_PULL_POLICY in rook operator configmap which should be used to
customize the imagePullPolicy for the csi driver and imagePullPolicy property in cephVersionSpec for ceph pods.
Signed-off-by: Avan Thakkar <athakkar@redhat.com>
if Multus is enabled the clusterinfo should be updated with
network as multus as to run the ceph cmds in remote
executor
Signed-off-by: parth-gr <paarora@redhat.com>
stability issues have been observed with 1s.
socket latency is expected whenever CPUs are
under minor pressure. Increasing value to
5s should cover most small-medium scale envs.
Resolves BZ: 2126566
Signed-off-by: Randy J. Martinez <randy@cephtips.com>
we need use `!=` for string comparision in bash instead of
`-ne`. Also, need to correct periodicity if condition to
make it work as expected.
Signed-off-by: subhamkrai <srai@redhat.com>
this commits add new field `MaxLogSize` inside `LogCollectorSpec` of
cephCluster cr which will take max size of log after which we want to
rotate the log.
Signed-off-by: subhamkrai <srai@redhat.com>
With octopus coming to end of life, we remove support from
Rook for deploying Ceph Octopus and assume a min version of
Pacific v16. Any checks for octopus or earlier are removed
from the reconciles since they are obsolete.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
we have noticed multiple failures because of the probe
failing, most of the time it's due to fewer resources.
But increasing timeout fixes that, so increasing the
probe timeout to 2s from default 1s so that it will
give more time to probe before failing.
Signed-off-by: subhamkrai <srai@redhat.com>
The startup probe for the OSD has been too aggressive to kill the OSD
in case the OSD is taking some time to start. The OSD may be self-optimizing,
scrubbing, or some other internal operation before it is ready to start.
Rather than disable the startup probe completely, the default is now
to retry for two hours in case the OSD is performing those operations.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The CSI package needs to load clusterInfo, today this code is in the mon
package which makes the call of LoadClusterInfo impossible without
having a circular import.
Signed-off-by: Sébastien Han <seb@redhat.com>
This removes double package imports. Example:
```
"github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
cephv1 "github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
```
Only one is now being used as shown in go-staticcheck ST1019
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
This introduces a new CRD to add the ability
to create rados namespace for a given
ceph block pool. Typically the name of the pool
is the name of the blockpool created by rook.
Closes: #7035
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Previously, the struct maintaining the list of cluster was still
initialized with a cluster item. Then the monitoring check will see that
the cluster is part of the struct already and thus won't run the
monitoring go routine again.
Now each time we cancel the context, we also remove the cluster item
from the map so that when the controller runs again, the monitoring
struct is re-populated and the go routine runs and statuses are updated.
Closes: https://github.com/rook/rook/issues/9911
Signed-off-by: Sébastien Han <seb@redhat.com>
The prometheus rules had been previously created if the cephcluster CR
setting monitoring.enabled was set to true. The rules were not customizable
and therefore not flexible enough. Now the rules are installed by the helm
chart. To customize the rules, a post-processor can be applied to the helm
chart.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
adding observedGeneration field in the cephcluster cr
status for having better control on reconciling,
as observedGeneration field will be updated by the controller
Closes: https://github.com/rook/rook/issues/9673
Signed-off-by: parth-gr <paarora@redhat.com>
Prior to this, we were comparing a pointer (the memory address) with a
struct. This was obviously always failing and returned false. We must
dereference the pointer to access the data contained at that memory
location.
Closes: https://github.com/rook/rook/issues/9544
Signed-off-by: Sébastien Han <seb@redhat.com>
Previously, we were ignoring configuration changes coming from the
operator's pod env variables. We were assuming most users were using the
operator configmap "rook-ceph-operator-config" but most Helm users
don't. Now we will reconcile if a cephcluster is found during a CREATE
event (can be an operator restart or a cephcluster creation) AND no
"rook-ceph-operator-config" is found which means the operator's pod env
var are used.
Also, unit tests have been added (long due) for the predicate!
Closes: #9602, #9487, #9579
Signed-off-by: Sébastien Han <seb@redhat.com>
Allow specifying daemon startup probes where we also allow configuring
liveness probes. Startup probes allow Rook to tolerate when Ceph daemons
occasionally take a long time to start up while not also making
Kubernetes liveness probes slower to detect runtime failures of daemons.
Startup probes are beta in Kubernetes 1.18, so we should not enable
probes by default for earlier Kubernetes versions.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This introduces a new CRD to add the ability to create subvolumegroup
for a given ceph filesystem volume. Typically the name of the volume is
the name of the filesystem created by rook.
Closes: https://github.com/rook/rook/issues/7036
Signed-off-by: Sébastien Han <seb@redhat.com>
Rook does not support running multiple clusters in the same namespace,
so the operator should not reconcile if a new cluster is added.
The scenario where a cluster is added while the operator is down is also
handled. CR updates are also handled.
When the operator detects more than one cluster it will refuse to
reconcile the CephCluster, and child CRDs will block too until the
operator is ready.
The user must remove one of the clusters before can continue to perform
any reconcile.
Closes: https://github.com/rook/rook/issues/9452
Signed-off-by: Sébastien Han <seb@redhat.com>
The built-in pools device_health_metrics and .nfs created by ceph
need to be configured for replicas, failure domain, etc.
To support this, we allow the pool to be created as a CR.
Since K8s does not support underscores in the resource names
the operator must translate this special pool name into
the name expected by ceph.
This also sets the basis for allowing filesystem data
pools to specify the desired pool name instead of requiring
a generated name.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
For now, we must run the container with UID 0 and privileged for
multiple reasons:
* the rook binary writes ceph config to /var/lib/rook which is owned by
root
* it's difficult to use /etc/ceph since it will conflict with the
rook-ceph-override configmap AND is also owned by root since it's a
mounted configmap.
* using /etc/ceph might be possible but has other issues with rook's
exec package since the ceph config is built from /var/lib/rook
Closes: https://github.com/rook/rook/issues/9385
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit adds context parameter to k8sutil replicaset and secret
functions. By this, we can handle cancellation during API call of
replicaset and secret resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil endpoint functions. By
this, we can handle cancellation during API call of endpoint resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>