Commit Graph
3241 Commits
Author SHA1 Message Date
Travis Nielsen 9d2aa1f6bd test: generate long node name depending on test suite
The generation of a long node name in the integration tests was
being done based on the k8s version. In the past, older K8s versions
did not support the changing name. Now it's more maintainable if
we generate the long name depending on the test suite.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-06 08:03:56 -07:00
Travis Nielsen ee83ca74af osd: truncate osd prepare job names further
In K8s 1.22 there is a bug in the job name generation that
the job name is truncated an additional 10 characters. This can cause an issue
in the generated pod name if it then ends in a non-alphanumeric character. In that case,
we more aggressively generate a hashed job name.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-06 08:02:04 -07:00
Travis Nielsen d55669293a mon: set stretch tiebreaker reliably during failover
The failover of the arbiter mon in a stretch cluster was sometimes
failing due to the new tiebreaker not being set in ceph.
Rook would repeatedly try to remove the old tiebreaker mon
and keep failing because the new tiebreaker had not been set.
Now we make setting the tiebreaker idempotent in case the operator
restarts in the middle of the operation or some other corner
case causes the expected tiebreaker to be set. In that case,
the next reconcile will also ensure the tiebreaker mon is
set as expected.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-02 12:50:07 -07:00
Travis Nielsen 32a884ac18 osd: honor skipUpgradeChecks for osds
Skipping upgrade checks was not being honored for OSDs.
Now the flag will be checked and allow the OSDs to be upgraded
without checking for the ok-to-stop condition.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-12-01 14:41:53 -07:00
Humble Chirammal e5f5d9be9f csi: mount host's /etc/selinux in node plugins
This commit introduces a new configuration option for
ceph csi driver to enable hostpath mounting of /etc/selinux
directory from the cluster node where csi plugin pods are
running, which inturn help the csi driver to specify
selinux-related mount options like context.

Ref# https://github.com/ceph/ceph-csi/issues/2295

The default value for this configuration is true and if cluster
nodes are running without selinux enabled, an admin can deploy
csi pods by specifying this option to `false` which skip the
host path mounting for the csi pods.

Signed-off-by: Humble Chirammal <hchiramm@redhat.com>
2021-12-01 11:13:26 +05:30
Sébastien Han 4a5ba05777 Merge pull request #9272 from olivierbouffet/fix9100
object: fix search user in objectstore
2021-11-30 16:23:30 +01:00
Sébastien Han decc481d40 Merge pull request #8740 from leseb/examples-layout
core: change directory layout
2021-11-30 16:02:37 +01:00
Sébastien Han c890710b63 core: change directory layout
As per discussion, proposing a new layout for the charts/yaml/olm files.

./deploy
├── charts
│   ├── rook-ceph
│   │   └── templates
│   └── rook-ceph-cluster
│       └── templates
├── examples
│   ├── csi
│   │   ├── cephfs
│   │   └── rbd
│   ├── flex
│   ├── monitoring
│   ├── pre-k8s-1.16
└── olm
    └── assemble

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-30 09:12:53 +01:00
Olivier Bouffet 5edeff4df8 object: fix search user in objectstore
avoid failed reconcile when multiple multisite objectstore are configured

Signed-off-by: Olivier Bouffet <olivier.bouffet@infomaniak.com>
2021-11-29 20:38:22 +01:00
Sébastien Han 83f7c2be87 Merge pull request #9259 from parth-gr/deviceClass
osd: update existing OSDs with deviceClass
2021-11-29 18:06:08 +01:00
parth-gr cbe505d122 osd: update existing OSDs with deviceClass
If we apply useAllNodes to false for the current deployment,
the OSDs should get updated with the individual nodes values and config,
The deviceClass was not updating to the existing OSDs because there was
bug in the check.
The check osdInfo.DeviceClass == "" which should be
checked like this osdInfo.DeviceClass == "None"

Updated the code so OSDs can make use of the devices present

Signed-off-by: parth-gr <paarora@redhat.com>
2021-11-29 21:40:02 +05:30
Sébastien Han bf662ec7b9 Merge pull request #9264 from leseb/fix-9151
cephfs-mirror: various fixes for random bootstrap peer import errors
2021-11-29 15:45:14 +01:00
Sébastien Han bd962e3e85 Merge pull request #9104 from BlaineEXE/nfs-restart-with-configmap
nfs: restart nfs servers when configmap is updated
2021-11-26 17:06:16 +01:00
Sébastien Han d8a8b05c1f cephfs-mirror: use combined output to possibly catch peer import error
By adding a combined output to the executor we might be able to fetch more
error messages.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-26 15:53:30 +01:00
Sébastien Han 7f7a72d941 core: add the ability to execute ceph commands with a combined output
Sometimes Ceph uses a different standard output to return errors or
merges standard error to standard out. So let's allow some commands to
return both in the output.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-26 15:52:49 +01:00
Sébastien Han d1cdba420a cephfs-mirror: try to mitigate peer import error
Some users have reported issues while adding the token, this is not
always reproducable so perhaps it's a typo when importing the token and
adding trailing spaces.

Closes: https://github.com/rook/rook/issues/9151
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-26 15:45:01 +01:00
Olivier 561cede1ef object: fix rgw ceph config
use Zone and ZoneGroup instead of storename for rgw_zone and rgw_zonegroup

Signed-off-by: Olivier Bouffet <olivier.bouffet@infomaniak.com>
(cherry picked from commit c92270cd66)
2021-11-25 18:38:42 +01:00
Sébastien Han 7402c2cce6 osd: check if osd is ok-to-stop before removal
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.

So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-25 10:21:32 +01:00
Sébastien Han 3e299fbe13 Merge pull request #9236 from leseb/fix-9234
core: fix openshift security context
2021-11-24 18:27:29 +01:00
Mara Sophie Grosch 733878b1eb monitoring: update label on prometheus resources
Updating the promethes reources (PrometheusRule and ServiceMonitor) is
done by fetching the current resource from the server and updating the
spec on it. This commit makes it also apply the labels, so users can
update them via rook CRDs.

Closes: https://github.com/rook/rook/issues/9241
Signed-off-by: Mara Sophie Grosch <littlefox@lf-net.org>
2021-11-24 17:17:02 +01:00
Sébastien Han e008464327 Merge pull request #9224 from leseb/fix-9205
nfs: only set the pool size when it exists and always run default pool creation
2021-11-24 14:36:14 +01:00
Sébastien Han f3142847d3 nfs: always run default pool creation
Previously, if the pool was present we would not run the pool creation
again. This is a problem if the pool spec changes, the new settings will
never be applied.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-24 14:04:19 +01:00
Sébastien Han b38f430c26 core: fix openshift security context
The MKNOD capability was missing and due to recent addition some pod now
only require this cap as well as privileged.
The cap must be explicitly exposed so it can be requested by a pod.

Closes: https://github.com/rook/rook/issues/9234
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-24 11:22:47 +01:00
Travis Nielsen 801b5d65a7 Merge pull request #9180 from LittleFox94/8502-monitoring-labels
#8502: Make monitoring object labels overridable
2021-11-23 09:44:06 -07:00
Sébastien Han c3accdcc3a nfs: only set the pool size when it exists
For CRD not using the new nfs spec that includes the pool settings,
applying the "size" property won't work since it is set to 0. The pool
still gets created but returns an error. The loop is re-queued but on
the second run the pool is detected so no further configuration is done.

Closes: https://github.com/rook/rook/issues/9205
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-23 11:07:43 +01:00
Sébastien Han e9f9a40138 Merge pull request #9212 from travisn/admin-test-cluster
core: Ensure cluster name is available on cluster info
2021-11-22 11:50:55 +01:00
Blaine Gardner 03ba7dec64 pool: file: object: clean up stop health checkers
Clean up the code used to stop health checkers for all controllers
(pool, file, object). Health checkers should now be stopped when
removing the finalizer for a forced deletion when the CephCluster does
not exist. This prevents leaking a running health checker for a resource
that is going to be imminently removed.

Also tidy the health checker stopping code so that it is similar for all
3 controllers. Of note, the object controller now uses namespace and
name for the object health checker, which would create a problem for
users who create a CephObjectStore with the same name in different
namespaces.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-19 10:29:12 -07:00
Travis Nielsen 1afd322650 core: ensure cluster name is available on cluster info
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.

Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-18 15:30:26 -07:00
Blaine Gardner cd832e5832 Merge pull request #8579 from thotz/vaultsslsupport
ceph: add support in RGW to communicate vault with TLS
2021-11-18 12:32:38 -07:00
parth-gr aea2856566 mon: set cluster name to mon cluster
the mon cluster clusterInfo is intiated seprately,
and misses out to set the cluster name and use default name as testing
from AdminClusterInfo.

Part-of: https://github.com/rook/rook/issues/9159
Signed-off-by: parth-gr <paarora@redhat.com>
2021-11-18 20:28:51 +05:30
Mara Sophie Grosch 00a5debbab monitoring: allow overriding monitoring labels
The templates for the mgr-generated ServiceMonitor and PrometheusRule objects included the labels
prometheus and team, making it impossible to override them as user.

This adds a new method `OverwriteApplyToObjectMeta` to
pkg/apis/ceph.rook.io/v1.Labels, which, contrary to the existing
`ApplyToObjectMeta` method, overwrites existing labels.

Closes: https://github.com/rook/rook/issues/8502
Signed-off-by: Mara Sophie Grosch <littlefox@lf-net.org>
2021-11-17 17:32:05 +01:00
Omar Pakker 8f9055809f osd: add privileged support (back) to blkdevmapper securityContext (work-around)
The blockdevmapper securityContext was changed to request a minimal set of
required capabilities for its operation and drop running as privileged.
While the base change works and is valid in terms of the container's copy operation,
it turns out that OpenShift may require some additional configuration not
currently covered by the limited securityContext and the capabilities granted.

To not break those OpenShift deployments, make the blkdevmapper securityContext
listen to the ROOK_HOSTPATH_REQUIRES_PRIVILEGED flag again to set privileged mode.
This flag is true on OpenShift deployments and running as privileged
works around the (missing) configuration problem for now.
To properly drop privileged completely some additional investigation needs
to be done on OpenShift deployments without relying on privileged execution.

Signed-off-by: Omar Pakker <Omar007@users.noreply.github.com>
2021-11-17 12:25:13 +01:00
Jiffin Tony Thottan aba50d3ca9 object: add support in RGW to communicate vault with TLS
From ceph v16.2.6 onwards the vault TLS suppport in RGW was added,
include similar changes for RGW.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-11-17 10:19:28 +05:30
Travis Nielsen 54a2b56b0c Merge pull request #9167 from parth-gr/mons-CRname
mon: update mons Cluster ClusterInfo with CR name
2021-11-15 11:11:09 -07:00
Sébastien Han c5783a77cf Merge pull request #9163 from y1r/add-context-k8sutil-node
core: add context parameter to k8sutil node
2021-11-15 16:25:13 +01:00
Travis Nielsen 716baf7785 Merge pull request #9158 from Omar007/fix/blkdevmapper-capabilities
osd: set blkdevmapper capabilities
2021-11-15 08:05:12 -07:00
Sébastien Han d5643b439d Merge pull request #9161 from y1r/add-context-k8sutil-daemonset
core: add context parameter to k8sutil daemonset
2021-11-15 15:14:44 +01:00
Sébastien Han f44b943a7c Merge pull request #9162 from y1r/add-context-k8sutil-job
core: add context parameter to k8sutil job
2021-11-15 15:13:54 +01:00
Satoru Takeuchi 07e1ca678a Merge pull request #9168 from cybozu-go/core-fix-unnecessary-option
core: remove unnecessary option
2021-11-15 22:50:08 +09:00
Yuichiro Ueno 4cc716a7ca core: add context parameter to k8sutil node
This commit adds context parameter to k8sutil node functions. By this,
we can handle cancellation during API call of node resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:45:59 +09:00
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
parth-gr 242f98e94b mon: update mons Cluster ClusterInfo with CR name
the mons re-initialize its ClusterInfo which results in missing
CR name on the ClusterInfo
Set the CR name to the mons cluster ClusterInfo

Closes: https://github.com/rook/rook/issues/9159
Signed-off-by: parth-gr <paarora@redhat.com>
2021-11-15 18:52:55 +05:30
Omar Pakker 4726d39688 osd: set blkdevmapper capabilities
The OSD blkdevmapper init container relies on the MKNOD capability,
which it does not actually request.
As a result, deployments fail on Kubernetes clusters that do not
happen to assign this capability to all containers by default.
Solve this by updating the container spec securityContext to
explicitly request the capability it relies on.

Closes: https://github.com/rook/rook/issues/9156
Signed-off-by: Omar Pakker <Omar007@users.noreply.github.com>
2021-11-15 13:11:19 +01:00
Satoru Takeuchi 77e7c249cc core: remove unnecessary option
rook command doesn't interpret `logtostderr` option. It's OK to just
remove this option because `capnslog` outputs all logs to stdout
by default. It's better to keep `AddGoFlagSet()` call because
some libraries might define their own flags with Go's `flag` package.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-11-15 11:53:45 +00:00
Sébastien Han ecd7fa7880 Merge pull request #9164 from y1r/add-context-k8sutil-pod
core: add context parameter to k8sutil pod
2021-11-15 11:58:36 +01:00
Yuichiro Ueno 0559977b8a core: add context parameter to k8sutil pod
This commit adds context parameter to k8sutil pod functions. By this, we
can handle cancellation during API call of pod resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 15:39:41 +09:00
Yuichiro Ueno e278812cc1 core: add context parameter to k8sutil daemonset
This commit adds context parameter to k8sutil daemonset functions. By
this, we can handle cancellation during API call of daemonset resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 15:09:31 +09:00
Yuichiro Ueno 0b575703c7 core: add context parameter to k8sutil deployment
This commit adds context parameter to k8sutil deployment functions. By
this, we can handle cancellation during API call of deployment resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-13 14:58:13 +09:00
Travis Nielsen 9ecd0cbc75 rgw: raise errors when rgw daemon fails to be created
The generation of the rgw deployment spec was swallowing errors
if any issues are raised such as the tls cert not being found
as expected in some configurations. We need to fail the reconcile
so the error will be logged and the admin can identify the issue.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-11-10 09:13:03 -07:00
Travis Nielsen 54d532b623 Merge pull request #9125 from sp98/update-scc
core: add `AllowHostDirVolumePlugin: true` to SCC and remove unused imports
2021-11-09 07:26:28 -07:00