Commit Graph
3023 Commits
Author SHA1 Message Date
Travis Nielsen 65f413ce1f Merge pull request #12286 from subhamkrai/fix-node-loss-rbd
core: faster recovery from rbd rwo node loss
2023-07-07 10:35:59 -06:00
subhamkrai 39b5c057ce core: faster recovery from rbd rwo node loss
in the existing node watcher, we'll check for node update
event and see if there are `out-of-service` taints are applied
and `ROOK_WATCH_FOR_NODE_FAILURE` is enabled in rook-ceph-operator-configmap,
if then we'll create the networkFence cr and delete the cr if nodes come back.
And, added the unit test too.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-07-07 21:16:25 +05:30
Travis Nielsen 3ffec84100 Merge pull request #12462 from Madhu-1/fix-holderpod
csi: update csi holder daemonset template
2023-07-07 08:32:00 -06:00
Travis Nielsen 18e2502a69 Merge pull request #12295 from subhamkrai/remve-default-scc
security: remove default scc privileges
2023-07-07 08:24:37 -06:00
subhamkrai 48f84b37e8 security: drop all capabilities to the container
adding drop `ALL` capabilities in rook operator container
as this is not required and will remove warning in ocp cluster.

Signed-off-by: subhamkrai <srai@redhat.com>
2023-07-07 17:06:48 +05:30
Travis Nielsen 46dd374e30 Merge pull request #12478 from Rakshith-R/csi/rm-checks
csi: drop checks for k8s version below the support minimum version
2023-07-06 11:53:47 -06:00
Rakshith R aa0aa34c23 csi: update csi sidecars' image version
Signed-off-by: Rakshith R <rar@redhat.com>
2023-07-06 16:18:35 +05:30
Rakshith R 76f4b1f586 csi: remove code related to betav1CsiDriver
This commit removes code related to betav1CsiDriver
since minimum k8s version support by rook is now
k8s v1.22 which does not support betav1 csi driver crd.

Signed-off-by: Rakshith R <rar@redhat.com>
2023-07-06 15:55:57 +05:30
Rakshith R d6b8a8f380 csi: remove k8s version check for oidc token
This commit removes k8s version check for oidc token
which required k8s v1.20 or above since minimum
k8s version that rook supports now is v1.22.

Signed-off-by: Rakshith R <rar@redhat.com>
2023-07-06 15:44:45 +05:30
Madhu Rajanna 1ba3aa4d19 csi: update csi holder daemonset template
Currently the holder daemonset is never updated
which will leaves the images the daemonset also
not updated. we should update the daemonset
template but not restart the csi holder pods
which causes the CSI volume access problem,
set the updateStrategy to OnDelete (already set in
yaml files) which allow us to update the holder
daemonset but not restart/update the pods, when a
pod is deleted or node is rebooted
the new changes will take effect.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2023-07-03 14:13:38 +02:00
travisn 4ff18dadf4 object: remove obsolete bucket health checker removal
The bucket health checker was removed in 1.10. Now in 1.12
we no longer need this removal of the bucket health
checker since it will no longer exist to remove.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-28 16:07:51 -06:00
Rakshith R fa1be14bca csi: update csiaddons k8s-sidecar to v0.7.0
This commit updates k8s-sidecar to v0.7.0
and all links now point to v0.7.0 tag.

Signed-off-by: Rakshith R <rar@redhat.com>
2023-06-28 20:38:29 +05:30
Rakshith R 1af70aca60 csi: update cephcsi to v3.9.0
This commit updates cephcsi image to
v3.9.0 release and also removes support
for v3.7.x.

refer:
https://github.com/ceph/ceph-csi/releases/tag/v3.8.0

Signed-off-by: Rakshith R <rar@redhat.com>
2023-06-28 20:38:29 +05:30
Rakshith R c488693cb3 csi: add SeLinuxMount: true Option to csidriver obj
This feature is alpha and available from k8s 1.25.
This enabled faster mounting of volumes using
RWOP access pod.

refer:
https://kubernetes.io/blog/2023/04/18/kubernetes-1-27-\
efficient-selinux-relabeling-beta/

Signed-off-by: Rakshith R <rar@redhat.com>
2023-06-28 20:23:51 +05:30
travisn 96ff80f8b0 ci: add exporter to commit prefixes
Add the exporter to the list of commit prefixes

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-27 11:03:03 -06:00
travisn 417f05311f exporter: ignore failed deletion of service monitor
If the ceph version does not suppport the exporter, the operator
will attempt to delete the service monitor related to the exporter
to ensure it does not exist for a node. If the expected rbac does
not exist, this will cause unnecessary errors since there would anyway
be no service monitor to delete. So we ignore the error of deleting
the service monitor for the exporter.

The context is also passed to the exporter so a new kubeconfig does
not need to be initiated and cause unnecessary logging about
invalid options for the config.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-26 12:26:09 -06:00
travisn 557a3e06cc core: api updates for controller runtime v0.15
For the controller runtime v0.15 there are some breaking
changes to the api that need to be updated.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-06-22 10:33:28 -06:00
Travis Nielsen d7dcf58f7c Merge pull request #12406 from polyedre/confusing-message
Fix confusing successful message when reconciling CephObjectStoreUser
2023-06-20 13:12:55 -06:00
Travis Nielsen 9a8e93cee8 Merge pull request #12341 from thotz/add-tls-secret-ref-in-objectstoreuser-secret
object : add ssl ref in cephobjectstore user secret
2023-06-20 10:20:35 -06:00
Lucas Henry 87bc3dfcdc operator: remove confusing successful message when reconciling CephObjectStoreUser
When creating a CephObjectStoreUser with a value spec.store that refers to an
unexisting CephObjectStore, after the reconciliation loop the
CephObjectStoreUser is in the ReconcileFailed state. However, a
ReconcileSucceeded event is created with this message:

"successfully configured CephObjectStoreUser"

The success message results of the return value for the error which is currently
`nil`. Let's replace it with the error message.

Signed-off-by: Lucas Henry <polyedre@disroot.org>
2023-06-20 11:02:04 +02:00
YZ775 ed9da20d68 operator: add ceph image version label to PVC
This PR makes Rook operator to add a label on PVC that contains an image version of Ceph when creating an OSD.

Signed-off-by: YZ775 <yuzuki-mimura@cybozu.co.jp>
2023-06-19 08:19:46 +09:00
Sachin Prabhu 8705f2b37e nfs: set correct kerberos domain in /etc/idmapd.conf
We add a new field domainName to the Kerberos section. The field is used
to setup /etc/idmapd.conf with the domain name. This allows idmapper to
map to kerberos credential to the correct uid/gid.

We add Spec.Security.Kerberos.DomainName to the CRD

Signed-off-by: Sachin Prabhu <sprabhu@redhat.com>
2023-06-09 18:06:45 +01:00
Jiffin Tony Thottan 6ea24cb11b object: add ssl ref in cephobjectstore user secret
There is no reference for ssl in cephobjectstore Secret, so users won't
have much idea why tls secret need to used. Hence give reference
object stores tls secret ref in the Secret.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-06-09 12:42:35 +05:30
Tarun Gupta Akirala df0ce26923 core: typo in logs to print fullname of CephCluster
printing namespace along with name would make debugging easier

Signed-off-by: Tarun Gupta Akirala <takirala@users.noreply.github.com>
2023-06-07 14:42:00 -07:00
avanthakkar 541d091f9c core: use ROOK_CEPH_MON_HOST from config store in OSD pods too
Signed-off-by: avanthakkar <avanjohn@gmail.com>

Volume "ceph-daemons-sock-dir" is coming empty in case if dataDirHostPath,
which is the case for osd onPVC. Fix the volume creation by using the
ceph cluster spec dataDirHostPath, which allows to run socket commands
on osd containers.
2023-06-01 21:09:58 +05:30
avanthakkar 41a8294510 core: cleanup exporter resources
Cleanup exporter daemon even if ceph version is not supported along with other resources like
metrics service and service monitor.
Signed-off-by: avanthakkar <avanjohn@gmail.com>
2023-06-01 02:04:17 +05:30
Travis Nielsen 6417ed4047 Merge pull request #12256 from thotz/add-missing-caps-object-user
object: add missing caps for object store user
2023-05-30 16:37:18 -06:00
Travis Nielsen 351751ac72 Merge pull request #12302 from Javlopez/feature/11429-use-default-for-logging-to-stderr
core: use -default-* flags
2023-05-30 13:05:02 -06:00
Javier 02e17196f2 core: use -default-* flags
enable flags with --default prefix for --log-to-stderr, --mon-cluster-log-to-stderr, --err-to-stderr, and --log-stderr-prefix

Signed-off-by: Javier <sjavierlopez@gmail.com>
2023-05-30 11:52:26 -06:00
Travis Nielsen 44a986396e Merge pull request #12293 from ulagbulag/fix-rook_cluster-bug-in-service-monitor
mgr: update default rook_cluster of ServiceMonitor
2023-05-30 11:10:52 -06:00
Ho Kim 42caef169f mgr: update default rook_cluster of ServiceMonitor
When deploying a cluster with Helm Chart, if the namespace is not the default value `rook-ceph`, and if the monitoring feature is enabled, then the generated ServiceMonitor's `rook_cluster` selector now follows the namespace, not the hard-coded value `rook-ceph`.

Signed-off-by: Ho Kim <ho.kim@ulagbulag.io>
2023-05-30 23:58:50 +09:00
Jiffin Tony Thottan 1f45cfa581 object: use networkspec from clusterinfo spec while running radosgw-admin
The radosgw-admin command uses the network spec from ceph cluster spec
in object context but it is not filled properly in the object package.
But with PR 10898, network spec is available in clusterinfo which can
be used directly. Also removed cluserspec from object context.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-05-30 11:06:03 +05:30
Jiffin Tony Thottan ad0c000e6a object: add missing caps for object store user
Lot of new caps added to rgw users, reflecting same changes on the
object store user CRD.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2023-05-26 17:36:00 +05:30
Blaine Gardner 73c39ca5b4 Merge pull request #11598 from opencmit/fix-11592
nfs: adding RGW section in ganesha.conf
2023-05-18 21:46:51 +02:00
Travis Nielsen 6cdf60a41d Merge pull request #12247 from travisn/rbd-reconcile-retry
rbdmirror: Retry reconcile if cluster not initialized
2023-05-17 17:25:54 -06:00
travisn 2877523137 rbdmirror: retry reconcile if cluster not initialized
The rbd mirror reconcile was not re-queuing the reconcile
if the cephcluster was not initialized. All other controllers
waiting for the initialization are requeuing the event,
just not the rbd mirror controller.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-05-17 15:47:17 -06:00
Travis Nielsen 9b99ab209c Merge pull request #12224 from travisn/dev-guide-review
docs: Dev guide review updates
2023-05-17 14:29:20 -06:00
Satoru Takeuchi 13e4464fed osd: supoprt expand lvm osd on pvc
The expansion of OSD on PVC is only supported in raw mode OSD.
It's nice to support lvm mode osd too.

Closes: https://github.com/rook/rook/issues/10835

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2023-05-16 00:23:09 +00:00
travisn 0f6b099403 docs: remove obsolete comments about old versions
Remove old mentions of rook or ceph versions that
are no longer relevant.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-05-11 11:45:04 -06:00
Travis Nielsen fd935f19a7 Merge pull request #12216 from travisn/no-service-mon-default
monitoring: Skip creating the service monitor for the exporter if monitoring is not enabled
2023-05-10 09:56:23 -06:00
travisn 1b80767877 monitoring: no service monitor by default for exporter
The exporter is enabled by default, but the service monitor
can only be enabled if prometheus CRDs are available. The
monitoring.enabled must be set to true as the flag that
prometheus is available and the service monitor should
be created.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-05-09 14:52:40 -06:00
avanthakkar 8a8eb863d0 core: add termination grace period for exporter pods
Shorten the termination time for exporter pod before it
gets deleted to avoid crashing of pod.
Signed-off-by: avanthakkar <avanjohn@gmail.com>
2023-05-10 02:16:12 +05:30
travisn 92a0ab37e3 monitoring: configurable option to disable prometheus metrics
The prometheus mgr module and ceph exporter can now be optionally
disabled by the monitoring.metricsDisabled setting in the
CephCluster CR. These will not be disabled by default, rather
than the mgr module being disabled by default from v1.11.4.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-05-05 16:04:51 -06:00
Madhu Rajanna 6f30700f5e core: disable controller runtime metrics server
As we are not using the controller runtime
metrics we dont even need to start the server
as its just uses extra resouces and its not
much useful.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2023-05-05 11:02:47 +02:00
Madhu Rajanna 1139cefc94 core: unset all encryption configuration before setting
if the configuration is already set in the ceph
database, using `ceph config assimilate-conf`
will add the key and value if its missing but
it wont update the value if the key is already
present, To fix this problem we can remove the
key and add all the configurations once again.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2023-05-03 18:15:45 +02:00
Blaine Gardner 6500c11d57 nfs: pare down config file volume sources
Pare down the available volume sources for NFS config files so that the
CRD isn't unnecessarily huge. This allows us to recommend
`kubectl apply` again in the upgrade doc.

Size of the NFS CRD is reduced by approximately 60%.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2023-05-02 11:34:24 -06:00
parth-gr ed61d25fc3 mgr: fix ceph dashboard login issue
Dashboard ac-user-create cmd was taking more time
then the usual ceph command to run,
So increased the timeout to run the cmd, and
now dashboard admin user is sucessfully created

Closes: https://github.com/rook/rook/issues/12113
Signed-off-by: parth-gr <paarora@redhat.com>
2023-05-02 19:43:53 +05:30
Travis Nielsen 844712a989 Merge pull request #12137 from travisn/active-mgr-default
mgr: Default to active mgr label
2023-04-25 10:31:12 -06:00
travisn db04b127b0 core: remove obsolete journal size config value
The journal size was only applicable to the filestore OSD
format which has not been supported by rook since v1.2.
Remove the remaining obsolete setting from the examples and
code.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-04-24 15:54:53 -06:00
travisn 1fa3c5ae18 mgr: default to active mgr label
The active and standby labels are updated on the mgr pods
by the mgr sidecar. In the case of a single mgr, there is
not sidecar, so the dashboard and other mgr services were
not available when there was a single mgr. Now the mgr
pod will default to the active mgr status label so that
the single mgr case will succeed. In the case of two mgrs,
the sidecar will immediately update the standby mgr to
remove the active status label.

Signed-off-by: travisn <tnielsen@redhat.com>
2023-04-24 14:47:15 -06:00