The generation of a long node name in the integration tests was
being done based on the k8s version. In the past, older K8s versions
did not support the changing name. Now it's more maintainable if
we generate the long name depending on the test suite.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In K8s 1.22 there is a bug in the job name generation that
the job name is truncated an additional 10 characters. This can cause an issue
in the generated pod name if it then ends in a non-alphanumeric character. In that case,
we more aggressively generate a hashed job name.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The failover of the arbiter mon in a stretch cluster was sometimes
failing due to the new tiebreaker not being set in ceph.
Rook would repeatedly try to remove the old tiebreaker mon
and keep failing because the new tiebreaker had not been set.
Now we make setting the tiebreaker idempotent in case the operator
restarts in the middle of the operation or some other corner
case causes the expected tiebreaker to be set. In that case,
the next reconcile will also ensure the tiebreaker mon is
set as expected.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Skipping upgrade checks was not being honored for OSDs.
Now the flag will be checked and allow the OSDs to be upgraded
without checking for the ok-to-stop condition.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit introduces a new configuration option for
ceph csi driver to enable hostpath mounting of /etc/selinux
directory from the cluster node where csi plugin pods are
running, which inturn help the csi driver to specify
selinux-related mount options like context.
Ref# https://github.com/ceph/ceph-csi/issues/2295
The default value for this configuration is true and if cluster
nodes are running without selinux enabled, an admin can deploy
csi pods by specifying this option to `false` which skip the
host path mounting for the csi pods.
Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
If we apply useAllNodes to false for the current deployment,
the OSDs should get updated with the individual nodes values and config,
The deviceClass was not updating to the existing OSDs because there was
bug in the check.
The check osdInfo.DeviceClass == "" which should be
checked like this osdInfo.DeviceClass == "None"
Updated the code so OSDs can make use of the devices present
Signed-off-by: parth-gr <paarora@redhat.com>
use Zone and ZoneGroup instead of storename for rgw_zone and rgw_zonegroup
Signed-off-by: Olivier Bouffet <olivier.bouffet@infomaniak.com>
(cherry picked from commit c92270cd66)
Updating the promethes reources (PrometheusRule and ServiceMonitor) is
done by fetching the current resource from the server and updating the
spec on it. This commit makes it also apply the labels, so users can
update them via rook CRDs.
Closes: https://github.com/rook/rook/issues/9241
Signed-off-by: Mara Sophie Grosch <littlefox@lf-net.org>
Previously, if the pool was present we would not run the pool creation
again. This is a problem if the pool spec changes, the new settings will
never be applied.
Signed-off-by: Sébastien Han <seb@redhat.com>
Clean up the code used to stop health checkers for all controllers
(pool, file, object). Health checkers should now be stopped when
removing the finalizer for a forced deletion when the CephCluster does
not exist. This prevents leaking a running health checker for a resource
that is going to be imminently removed.
Also tidy the health checker stopping code so that it is similar for all
3 controllers. Of note, the object controller now uses namespace and
name for the object health checker, which would create a problem for
users who create a CephObjectStore with the same name in different
namespaces.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
the mon cluster clusterInfo is intiated seprately,
and misses out to set the cluster name and use default name as testing
from AdminClusterInfo.
Part-of: https://github.com/rook/rook/issues/9159
Signed-off-by: parth-gr <paarora@redhat.com>
The templates for the mgr-generated ServiceMonitor and PrometheusRule objects included the labels
prometheus and team, making it impossible to override them as user.
This adds a new method `OverwriteApplyToObjectMeta` to
pkg/apis/ceph.rook.io/v1.Labels, which, contrary to the existing
`ApplyToObjectMeta` method, overwrites existing labels.
Closes: https://github.com/rook/rook/issues/8502
Signed-off-by: Mara Sophie Grosch <littlefox@lf-net.org>
The blockdevmapper securityContext was changed to request a minimal set of
required capabilities for its operation and drop running as privileged.
While the base change works and is valid in terms of the container's copy operation,
it turns out that OpenShift may require some additional configuration not
currently covered by the limited securityContext and the capabilities granted.
To not break those OpenShift deployments, make the blkdevmapper securityContext
listen to the ROOK_HOSTPATH_REQUIRES_PRIVILEGED flag again to set privileged mode.
This flag is true on OpenShift deployments and running as privileged
works around the (missing) configuration problem for now.
To properly drop privileged completely some additional investigation needs
to be done on OpenShift deployments without relying on privileged execution.
Signed-off-by: Omar Pakker <Omar007@users.noreply.github.com>
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
This commit adds context parameter to k8sutil node functions. By this,
we can handle cancellation during API call of node resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
the mons re-initialize its ClusterInfo which results in missing
CR name on the ClusterInfo
Set the CR name to the mons cluster ClusterInfo
Closes: https://github.com/rook/rook/issues/9159
Signed-off-by: parth-gr <paarora@redhat.com>
The OSD blkdevmapper init container relies on the MKNOD capability,
which it does not actually request.
As a result, deployments fail on Kubernetes clusters that do not
happen to assign this capability to all containers by default.
Solve this by updating the container spec securityContext to
explicitly request the capability it relies on.
Closes: https://github.com/rook/rook/issues/9156
Signed-off-by: Omar Pakker <Omar007@users.noreply.github.com>
This commit adds context parameter to k8sutil pod functions. By this, we
can handle cancellation during API call of pod resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil daemonset functions. By
this, we can handle cancellation during API call of daemonset resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
The generation of the rgw deployment spec was swallowing errors
if any issues are raised such as the tls cert not being found
as expected in some configurations. We need to fail the reconcile
so the error will be logged and the admin can identify the issue.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When the configuration configmap is updated for a CephNFS server, the
NFS application should restart to ensure it is running with the latest
config.
Fixes#9028
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
When a reconcile is started for OSDs, the prepare jobs are first
deleted from a previous reconcile. The timeout for the osd prepare
job deletion was only 40s. After that timeout, the reconcile attempts
to continue waiting for the pod, but of course will never complete
since the OSD prepare was not running in the first place, causing the
reconcile to wait indefinitely. In the reported issue, the osd prepare
jobs were actually deleted successfully, the timeout just wasn't long
enough. Pods need at least a minute to be forcefully deleted,
so we increase the timeout to 90s to give it some extra buffer.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The rook operator as well as the toolbox pod run with the "rook" user
with UID 2016. The UID was chosen based on the year of the initial
commit in the rook/rook repository.
No more root user running.
Closes: https://github.com/rook/rook/issues/8734
Signed-off-by: Sébastien Han <seb@redhat.com>
In the event a ceph image is specified that is lower than the current running
version of the daemons, the downgrade is allowed, even if not technically
supported. All of the core daemons (mon,mgr,osd) were being downgraded,
but the daemons for other controllers (rgw,mds,rbdmirror) were not being
downgraded, resulting in an inconsistent cluster. Now we log that the downgrade
is not supported and all all of the daemons to be downgraded.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
rgw doesn't respond `livenessProbe` if the number of connection reaches its
limit (by default, 1000). Then rgw is out of service but still live.
Hense the current `livenessProbe` logic is suitiable for `readinessProbe`.
`tcpSocket` is enough for `livenessProbe`.
Closes: #8407
Signed-off-by: Yuzuki Mimura <yuzuki725.m@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
On Pacific, Ceph's default is "upmap", so we should let it be
like this.
This lets the user change the mode is desired.
On Octopus though, Rook continues to force the mode to "upmap".
Closes: https://github.com/rook/rook/issues/9062
Signed-off-by: Sébastien Han <seb@redhat.com>
Ths NFS spec now supports the CephBlockPool spec which means that it can
take advantage of all the known settings like compression, size, failure
domain etc.
Closes: https://github.com/rook/rook/issues/9034
Signed-off-by: Sébastien Han <seb@redhat.com>