Commit Graph
121 Commits
Author SHA1 Message Date
Sébastien Han bd962e3e85 Merge pull request #9104 from BlaineEXE/nfs-restart-with-configmap
nfs: restart nfs servers when configmap is updated
2021-11-26 17:06:16 +01:00
Sébastien Han f3142847d3 nfs: always run default pool creation
Previously, if the pool was present we would not run the pool creation
again. This is a problem if the pool spec changes, the new settings will
never be applied.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-11-24 14:04:19 +01:00
Yuichiro Ueno 3799542356 core: add context parameter to k8sutil job
This commit adds context parameter to k8sutil job functions. By this, we
can handle cancellation during API call of job resource.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-11-15 22:39:08 +09:00
Blaine Gardner 5a7dee2f63 nfs: restart nfs servers when configmap is updated
When the configuration configmap is updated for a CephNFS server, the
NFS application should restart to ensure it is running with the latest
config.

Fixes #9028

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-11-08 11:01:09 -07:00
Sébastien Han e0145b9643 nfs: add pool setting CR option
Ths NFS spec now supports the CephBlockPool spec which means that it can
take advantage of all the known settings like compression, size, failure
domain etc.

Closes: https://github.com/rook/rook/issues/9034
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-10-29 09:04:31 +02:00
Travis Nielsen fd10d98dc6 core: treat cluster as not existing if the cleanup policy is set
The cluster CR can be forcefully deleted and cleanup the
cluster resources if the yes-really-destroy-data policy
is set on the CR. In this case, the other controllers should
treat the cluster CR as not existing and allow the finalizers
to be removed on those resources if they are requested for
deletion.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-10-27 10:25:06 -06:00
Blaine Gardner 2f850b6ae6 Merge pull request #8613 from subhamkrai/remove-nautilus
ceph: remove ceph nautilus, ceph octopus to default
2021-10-25 09:09:18 -06:00
Yuichiro Ueno 3fd86f83ae core: add context parameter to opcontroller
This commit adds context parameter to utilities in opcontroller to
remove context.TODO use in opcontroller. By this, we can handle
cancellation of reconcilers in a fine-grained way.

Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
2021-10-25 20:45:06 +09:00
subhamkrai 0150966024 ceph: remove ceph nautilus, ceph octopus to default
since rook 1.8, ceph nautilus no longer supported,
ceph octopus will be the minimum ceph version.

Closes: https://github.com/rook/rook/issues/7908
Signed-off-by: subhamkrai <srai@redhat.com>
2021-10-20 14:45:12 +05:30
Joseph Sawaya ee791b09fd ceph: update CephNFS to use ".nfs" pool in newer ceph versions
This commit updates the CephNFS CR to make the RADOS settings optional
for Ceph versions above 16.2.7 due to the NFS module changes in Ceph.
The changes in Ceph make it so the RADOS pool is always ".nfs" and the
RADOS namespace is always the name of the NFS cluster.

This commit also handles the changes in Ceph Pacific versions before 16.2.7
where the default pool name is "nfs-ganesha" instead of ".nfs".

Closes: https://github.com/rook/rook/issues/8450
Signed-off-by: Joseph Sawaya <jsawaya@redhat.com>
2021-10-04 17:39:47 +00:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Sébastien Han 2d55e69416 ceph: move scheme initialization to the same place
Let's initialize the schemes in a single place instead of doing it
when each controller initializes.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-08-06 11:33:10 +02:00
Sébastien Han d55fb8d13b Merge pull request #8351 from leseb/refact-exec-helpers
ceph: remove unnecessary exec helpers
2021-07-23 18:12:57 +02:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Sébastien Han 6bce1ff3e9 ceph: move all of our docker.io reference to quay.io
Recently, the builds of `ceph/ceph` image moved to quay.io, see
https://github.com/ceph/ceph-build/pull/1883 for more details.
Current images will remain but new builds will happen on quay.io only.

This means that tags such as `v14.2`, `v15.2`,`v16.2` will need to
switch to quay.io to get updates.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-22 11:17:05 +02:00
Travis Nielsen 4b1cf51ef1 ceph: scaling down nfs deployment was failing
When scaling down the nfs deployment gracefully, the wrong error code was checked,
resulting in a failed scaling down even though the condition was expected.
Now we check properly if the deployment was removed properly.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-07-02 11:52:03 -06:00
Blaine Gardner c22f545ebf ceph: block delete object store when buckets exist
Block deletion of CephObjectStore resources when buckets exist in the
object store.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-06-29 14:31:39 -06:00
Sébastien Han 90bea8a560 ceph: stop using radosgw-admin CLI for s3 user management
We have been having many issues with external mode with Ceph version
mismatching. The operator would have a Ceph version different than the
external cluster. The `radosgw-admin` was used to interact with S3
users, even a small version delta would cause the command to coredump.
After checking with the rgw core team it appears Rook was misusing the
CLI and the admin ops API should be used instead.
So this patch is the first introduction of go-ceph in Rook to consume
the rgw admin ops API instead of the `radosgw-admin` CLI, **only** for
user management in this initial commit.
Later we can do more such as bucket operation, zone management etc.

Closes: https://github.com/rook/rook/issues/7924
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-09 11:08:22 +02:00
Travis Nielsen b0a63711f5 build: refactor to consolidate the rook.io/v1 package
The rook.io/v1 package was only an internal implementation detail and
does not have any CRDs that rely on it. The CRD deserialization should
handle the change in internal types without any issue. This separation
gives more flexibility for the storage providers to implement exactly
what is needed for their storage provider instead of forcing to use the
same types and risk affecting another storage provider.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-05-18 19:55:37 -06:00
Lars Lehtonen 41c567beaf ceph: fix multiple imports
This fixes additional double-imports.

Signed-off-by: Lars Lehtonen <lars.lehtonen@gmail.com>
2021-04-13 02:02:09 -07:00
Travis Nielsen 721acd1a8a Revert "ceph: added rook-ceph-default service account"
This reverts commit 737fb099fe.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-04-08 11:46:12 -06:00
parth-grandTareq Sharafy 737fb099fe ceph: added rook-ceph-default service account
When a private docker registry is used and an image pull secret is specified in the chart, the pods with default Service Account fail to pull the image due to authentication issues.
Added rook-ceph-default service account and modify the pods specifications by adding the serviceAccountName.

Closes: https://github.com/rook/rook/issues/6673
Co-authored-by: Tareq Sharafy <tareq.sha@gmail.com>
Signed-off-by: parth-gr <partharora1010@gmail.com>
2021-03-29 20:20:40 +05:30
Satoru Takeuchi 9eb3160d8b ceph: validate all owner references
Remaining work of #7259

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-26 06:43:54 +00:00
subhamkraiandTravis Nielsen 319e4a41a4 ceph: placement in case of both PVC and non-PVC's
In the case of PVC,
We are giving lower priority to all placement.
We want deviceSet placement to applied and
override in case of overlapping settings and
we are merging nodeAffinity if applied in both
all placement and deviceSet.

In case of non-PVC,
we apply spec.placement

Signed-off-by: subhamkrai <srai@redhat.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
2021-03-18 10:32:57 +05:30
Satoru Takeuchi 26c8fd9bd1 ceph: improve owner reference management
It's better to validate ownerReferences when setting them. In addition, we should use
controllerrutil.Set{Controller,Owner}Reference, that have such validation, as possible.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-16 10:29:33 +00:00
Sébastien Han 86de4f6fb3 Merge pull request #7238 from varshar16/rados-namespace-optional
ceph: rados pool namespace is optional for nfs deployments
2021-02-17 11:40:06 +01:00
Varsha Rao 7a4ebbef0e ceph: rados pool namespace is optional for nfs deployments
There is no need to validate rados pool namespace, it is optional[1]. And if in
case rados namespace is not specified, it is appropriately handled.

[1] https://github.com/nfs-ganesha/nfs-ganesha/blob/next/src/doc/man/ganesha-core-config.rst

Signed-off-by: Varsha Rao <varao@redhat.com>
2021-02-16 16:25:50 +05:30
subhamkrai 6158b6608c ceph: do not merge nodeAffinity
we don't want to merge the node affinity if specified in both
storageClassDeviceSet and placement. Now, we are passing new
boolean argument in `ApplyToPodSpec` which will be false in
case of osd and prepare pods because we don't want to overlap
placement of osd on PVC's and non-PVC's.

Signed-off-by: subhamkrai <srai@redhat.com>
2021-02-16 12:25:28 +05:30
Varsha Rao baa79cf4c3 ceph: filesystem is not required for nfs server deployment
nfs server requires valid pool for storing nfs-ganesha recovery objects. Only
when filesystem exports are created, filesystem should be configured first.

Signed-off-by: Varsha Rao <varao@redhat.com>
2021-02-08 20:13:14 +05:30
Sébastien Han 8d033efb5a ceph: silence harmless errors
If the ceph cli outputs an error with "error calling conf_read_file"
this means that the operator has not written its ceph configuration
file. Thus ceph cli commands will fail, so we can just ignore that since
the operator will soon write this file in its initialization sequence.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-02-03 17:28:44 +01:00
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
Varsha Rao 9530807179 ceph: remove 'ganesha-' from nfs-ganesha config object name
PR[1] in volume nfs plugin will remove it too. Since it can cause
inconsistencies.

This change will not affect exports created using dashboard. As it looks for
ganesha config object just by 'conf-' [2].  Exports cannot be created by using
volume/nfs plugin in Octopus version. Because the ceph rook module is broken.

[1] https://github.com/ceph/ceph/pull/38510

[2] https://github.com/ceph/ceph/blob/octopus/src/pybind/mgr/dashboard/services/ganesha.py#L1111

Signed-off-by: Varsha Rao <varao@redhat.com>
2020-12-18 17:18:52 +05:30
Sébastien Han c6a87203ca ceph: add log collector
We can now collect logs directly into a side-car container.
A new CRD spec has been added:

spec:
  logCollector:
    enabled: true
    periodicity: 24h

Every 24h we will rotate log files for each Ceph daemon.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-12-01 16:30:17 +01:00
Sébastien Han 97be23e374 ceph: apply finalizer before updating object status
We must add the finalizer right after the object creation otherwise the
seerver will later return an error on update that the object has been
modified. Indeed, it has been by the task that updates the status when
the object is first created.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-18 18:16:59 +01:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Blaine Gardner 5fe80660dd ceph: fix nfs daemons not updating
The NFS reconciler was not updating NFS daemons that already existed.
The controller would fail to update daemons if the number of active
daemons did not increase or decrease.

Make sure existing NFS daemons are updated to new versions with all
scale up and scale down events.

Resolves #6611

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2020-11-12 11:16:24 -07:00
Varsha Rao 8fadaeeb09 ceph: add 'ganesha-' to nfs-ganesha config object name
Since ceph octopus version, the expected nfs-ganesha config object name by
mgr/volumes/nfs[1] plugin is "conf-nfs.ganesha-<clustername>". For older
versions the name is "conf-<clustername>.<nodeid>".

[1] https://github.com/ceph/ceph/blob/master/src/pybind/mgr/volumes/fs/nfs.py#L648-L655

Signed-off-by: Varsha Rao <varao@redhat.com>
2020-11-12 18:09:46 +05:30
subhamkrai 7d1b5f7416 ceph: change nfs-ganesha log level
this commit allows changing the default
log level from nfs.yaml file by adding
new command.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-11-11 20:34:07 +05:30
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
Varsha Rao f5d125f8c9 ceph: update ganesha config object name
In cephadm and volume/nfs module, all ganesha daemons within a cluster watch
single config object. Each ganesha daemon in rook watches it's own config
object. This patch updates the ganesha config object name to maintain
uniformity.

Signed-off-by: Varsha Rao <varao@redhat.com>
2020-10-08 14:42:25 +05:30
Varsha Rao 8d3140c8a0 ceph: create user for each ganesha daemon
Ganesha daemons don't need to acess ceph cluster with admin capabilities.
Instead new user with restricted caps can be used.

Signed-off-by: Varsha Rao <varao@redhat.com>
2020-09-29 22:27:17 +05:30
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
Varsha Rao a0dab8faab ceph: remove NParts and Cache_Size from MDCACHE block
As setting them to small value affects the performance and they are not related
to metadata caching. https://review.gerrithub.io/c/ffilz/nfs-ganesha/+/495185

Signed-off-by: Varsha Rao <varao@redhat.com>
2020-09-02 16:41:15 +05:30
Alexander Trost a87c00ddee ceph: allow user to add labels to pods
This allows users to specify additional labels to be added to the components
created by the Rook Ceph operator.
Additionally the rbd mirrors are now getting their annotations and
labels set.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2020-08-25 19:38:51 +02:00
Travis Nielsen 06c21954e1 ceph: labels are immutable for match selectors
The labels cannot be updated on a deployment match selector.
In v1.4 a new label was added for ceph_daemon_type that is intended
to be on the pod labels, but cannot be applied to the match selectors.
Therefore, we suppress any new labels that are added to the daemons.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-08-04 16:47:00 -06:00
Blaine Gardner dcf8ae2526 ceph: rename pod labels function for more clarity
Rename PodLabels function to CephDaemonAppLabels for more clarity about
what the function's purpose is.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-30 11:16:30 -06:00
Blaine Gardner 6ceaef8610 ceph: add new "ceph_daemon_type" label to nfs pods
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-30 11:16:29 -06:00
subhamkrai acc4ed5df9 ceph: handling gosec errors that are not checked
this commit handles all the gosec g104
i.e audit errors not checked inside pkg.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-24 20:02:06 +05:30
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00