Commit Graph
232 Commits
Author SHA1 Message Date
Travis Nielsen 53ed11f15b build: remove the edgefs operator from rook
The EdgeFS operator has been deprecated for some time in Rook.
If the replacement is added back to Rook it can be completed
according to the new guidelines in the documentation.
https://rook.io/docs/rook/master/storage-providers.html

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2021-01-25 17:51:24 -07:00
Sébastien Han 7558d37420 core: bump to controller-runtime 0.7.0 version
Now using https://github.com/kubernetes-sigs/controller-runtime/releases/tag/v0.7.0

Closes: https://github.com/rook/rook/issues/6689
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-13 11:00:43 +01:00
subhamkrai 0146eb3805 ceph: update to latest Kubernetes version 1.20.0
updating to latest Kubernetes version 1.20.0 fix
security issues. In the current version, it allows
for the token leak in logs when logLevel >= 9.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-12-21 13:18:13 +05:30
binoue 2ecde7c9c6 ceph: remove deplicated env value
Currently, mgr deployment has two overlapping environment variables, ROOK_POD_IP.
This causes kube-apiserver to create error logs

Signed-off-by: binoue <banji-inoue@cybozu.co.jp>
2020-12-15 09:11:26 +09:00
Travis Nielsen d0ba899341 ceph: check stretch cluster is ready to configure arbiter
The arbiter can only be configured with the stretch cluster if the
CRUSH map is balanced and there are two zones in the CRUSH map.
After the OSDs are configured, we wait for all the OSD pods to be
running and that the CRUSH map is balanced. If it takes more than
two minutes, we fail the reconcile and try again. This is only done
the first time the stretch cluster is configured. In future reconciles
we first check if the stretch cluster is already enabled before
enabling it again.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-09 16:37:24 -07:00
Travis Nielsen 156774c459 ceph: remove obsolete topology comments and fix comment typo
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-04 09:10:52 -07:00
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Arun Kumar Mohan 421f340c1c ceph: updating the dependencies for operator SDK v1.0.0
Updating the dependencies' versions to match with the newer Operator SDK
version v1.x

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:03:40 +05:30
Alexander Trost 34930020ab ceph: allow pod labels to be added to CSI components
This adds the functionality to add custom pod labels to the CSI
components through the operator configuration way of env vars or config
map.

Resolves #6593

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2020-11-12 13:06:24 +01:00
Satoru Takeuchi c7990777d2 ceph: make the placement of admission controller configurable
If users want to restrict the nodes where Ceph daemons should exist,
it's better to make the placement of admission controller configurable
as other daemons.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-11-09 11:54:52 +00:00
Sébastien Han e3d032dbc1 Merge pull request #6489 from travisn/stretch-mons
ceph: Stretch cluster configuration
2020-11-06 15:16:39 +01:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
Travis Nielsen 7c3cdbce50 ceph: use constants for k8s topology labels
Instead of directly defining our own constants, we should be using the k8s
constants for the well-known topology labels.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-05 09:23:58 -07:00
Travis Nielsen b91f4211c9 ceph: configure a stretched cluster
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-05 09:17:52 -07:00
subhamkrai 0fddfcf307 ceph: handle golangci-lint linter errcheck error
this commit handle golangci-lint linter errcheck.

`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases

To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-30 22:24:34 +05:30
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
Sébastien Han d118bd38a6 Merge pull request #6278 from subhamkrai/golanci-lint-unused
ceph: handle golangci-lint linter unused
2020-09-18 12:44:37 +02:00
Travis Nielsen 8664cbb8dd ceph: if spec diff checking fails during upgrade, assume it changed
If the pod spec changed, we expect an upgrade to proceed for that daemon.
If the check for a changed pod spec fails, we were skipping the update
of that daemon. Instead of skipping the update, we now assume the pod
spec changed if we fail to detect the change so that we can ensure the upgrade
even if we check for upgrades too often.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-09-17 13:22:53 -06:00
subhamkrai f9fafe62d4 ceph: handle golangci-lint linter unused
this commit will enable one more linter
in golangci-lint.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 22:53:24 +05:30
subhamkrai 25c116a4bd ceph: handle golangci-lint linter gosimple
this commit handles all the errors  check for
golangci-lint linter gosimple.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 15:10:27 +05:30
rohan47 8c4ede1edc ceph: added support for multus for csi
CSI pods now utilize multus networking and connect to public
network specified in the CephCluster CR.

Closes: https://github.com/rook/rook/issues/5356
Signed-off-by: rohan47 <rohgupta@redhat.com>
2020-08-20 18:25:09 +05:30
rohan47 f5bbc8c5c5 ceph: fix networkAttachmentConfig to be compatible with whereabouts cni
In the whereabouts config, the network range is determined by range
option unlike other cni plugins that use Subnet option.

Signed-off-by: rohan47 <rohgupta@redhat.com>
2020-08-05 21:17:13 +05:30
Travis Nielsen 704e80d1b5 Merge pull request #5912 from rohantmp/rohantmp-patch-1
Core: Add comment warning to keep k8sutil hashing consistent for name…
2020-07-28 10:41:47 -06:00
Rohan Joseph 82b062f91d core: add comment warning to keep k8sutil hashing consistent
The name generation utils need to provide consistent outputs across
versions so that names generated based on values are consistent and
older generated keys are not lost.

Signed-off-by: Rohan CJ <rohantmp@gmail.com>
2020-07-28 12:51:54 +05:30
subhamkrai 465a0f0aec ceph: handling gosec error code g601
this commit handles all the gosec g601
error code (i.e Implicit memory aliasing
of items from a range statement).

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-28 11:34:49 +05:30
Travis Nielsen 7bef54c089 edgefs: error checking for ineffectual assignments
All assignments made to errors or other variables should be
used instead of ignored. This checks for errors in a couple
places that were previously ignored, and also initializes
a variable such that it won't be always overwritten by the
various if else statements.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-20 17:11:37 -06:00
Travis Nielsen e74c7eaef8 ceph: refactor context and clusterInfo passed to the ceph commands
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.

Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:54 -06:00
Blaine Gardner 7117fc12b7 ceph: osd: add drive groups spec to cluster CR
Add the ability to provision Ceph OSDs with Drive Groups.
This adds Drive Groups to the CephCluster CRD, and it sets code
in place for propagating this config to the OSD provisioning pod.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2020-07-14 09:43:03 -06:00
subhamkrai 7f9f1690d2 ceph: remove csi drivers when disable
remove csi drivers and delete k8s
services that drivers create,when
the default setting of drivers changed
to false.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-06 13:33:52 +05:30
Ahmad Nurus S f7f8270f58 nfs: rewrite nfs operator controller to use controller-runtime
Signed-off-by: Ahmad Nurus S <prksu.sh@gmail.com>
2020-07-03 14:33:25 +07:00
Sébastien Han e4eaa91ede ceph: add rgw endpoint healthcheck
We have introduced a new goroutine to check the state of the rgw
endpoint. It will run every minute and perform operations on a bucket.
The success or failure will be reported as part of the status field of
the CephObjectStore CR.

A good status will look like:

status:
  endpointStatus:
    lastChanged: "2020-06-25T13:47:45Z"
    lastChecked: "2020-06-25T13:48:46Z"
  phase: Connected

A failed status:

status:
  endpointStatus:
    details: |-
      error creating bucket "rook-ceph-internal-s3-bucket-checker": RequestError: send request failed
      caused by: Put http://rook-ceph-rgw-my-store.rook-ceph:8080/rook-ceph-internal-s3-bucket-checker: dial tcp 10.108.189.148:8080: connect: connection refused
    health: ERROR

This check works for both converged and external modes. Note that the
CephObjectStore CRD has a new field called "externalRgwEndpoints" which
allows you to define a list of IP addresses pointing to rgws.

Closes: https://github.com/rook/rook/issues/5692
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-07-02 16:34:10 +02:00
Travis Nielsen ff6a3f80cd ceph: improve osd update logging
The operator should only print helpful info messages
when an OSD is going to be updated. If the OSD hasn't
changed it is just a debug message.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-28 10:05:51 -06:00
Travis Nielsen 171910707b ceph: refactor pod deletion to k8sutil package
The pod deletion is needed by other daemons besides the osds,
so we move the helper method into the k8sutil package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-28 10:05:50 -06:00
Madhu Rajanna a16d3bd422 k8sutil: Make client as the first argument
To have parity with other functions and
also client should be the first argument to
the function.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-05-14 11:02:14 +05:30
Travis Nielsen 65377cbe63 ceph: osd on pvc node assignment is node selector
An OSD on a PVC when portable=false is assigned to a node
with a node selector. The same node assignment is expected
for the lifetime of the cluster. On subsequent reconciles,
the operator was looking up the node assignment from the
nodeName on the pod spec, which is not set if the pod is down.
The operator needs to retrieve the assignment from the
deployment spec node selector.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-07 10:51:46 -06:00
Travis Nielsen cd806c83ee ceph: log the operator settings and their source
The operator settings can come either from the configmap, an
env var, or a default value. Knowing where the setting is loaded
from is an important detail to see in the log.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-05-05 13:10:01 -06:00
Travis Nielsen f8e884d2ce Merge pull request #5287 from Madhu-1/resource-limit-csi
CSI: Add resource request and limit for CSI pods
2020-04-29 08:09:19 -06:00
Madhu Rajanna 1bb5e1bab2 k8sUtil: Add YamlToContainerResource to resource
YamlToContainerResource can be used to convert the raw
string data to the array of ContainerResource.
the ContainerResource can be applied to the
pod container resources request and limits.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2020-04-28 13:13:31 +05:30
Sébastien Han f27fd207ce ceph: convert the CephCluster controller to the controller-runtime
This is the final conversion to controller-runtime conversion. This time the
CephCluster CRD has been converted to use the controller-runtime
library.
The controller incorporates all the previous watchers too, so the Node
and hot-plug configmap are been watched too.
Only the operator setting configmap is not being watcher since it's not
related to the CephCluster CRD.
Not only the patch converts to controller-runtime but also tries to
re-organize the tree of the repo to actually make the code more
readable and have better functions/methods/tests separations.

Closes: https://github.com/rook/rook/issues/4939
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-28 09:40:35 +02:00
Sébastien Han 20d1543507 ceph: add multus support
You can now use Rook along with Multus. Multus must be up and running
and the right ressources must exist such as NetworkAttachmentDefinition
CR.
The Cluster CR spec has new fields to work with multus:

network:
  provider: multus (or 'host' for hostNetworking)
  selectors:
    public: NetworkAttachmentDefinition name
    cluster: NetworkAttachmentDefinition name

If only a single NetworkAttachmentDefinition is provided Rook will use
both anyway for the Ceph traffic.

Please refer to the doc to learn more.

Closes: https://github.com/rook/rook/issues/4716
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-04-03 18:00:15 +02:00
Sébastien Han 9c343ae94e Merge pull request #5032 from umangachapagain/config-override
Ceph: add CSI configurations to ConfigMap
2020-03-27 15:21:24 +01:00
Umanga Chapagain 0e932c15eb Ceph: add CSI configurations to ConfigMap
This commit adds all the CSI configurations to ConfigMap.
This configMap can be used in combination with Env Vars
to configure Ceph CSI drivers in rook.

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-27 15:17:30 +05:30
Sébastien Han c62d89c7bc k8sutil: mitigate UpdateDeploymentAndWait condition
If a deployment stays in pending we should give early by looking at
ProgressDeadlineExceeded, this will reduce the time to wait from 20 min
to 10 min because ProgressDeadlineExceeded default is 600 seconds.

Prior to this patch we would wait 20min since we take
currentDeployment.Spec.ProgressDeadlineSeconds which is typically 600
then retry every 2 seconds, which makes it 20min total.

Closes: https://github.com/rook/rook/issues/5090
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-26 14:15:45 +01:00
Umanga Chapagain db3735656c Ceph: fix updates for servicemonitor
with older implementation, servicemonitor was not getting
updated due to missing resource version. This fix adds
resource version to the servicemonitor definition and
ensures that the object is properly created or
updated.

Signed-off-by: Umanga Chapagain <chapagainumanga@gmail.com>
2020-03-24 16:30:44 +05:30
Travis Nielsen 18b0e7d295 ceph: scrub ceph commands to write actions to the log
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
moricho 7d0215d820 ci: switch to use go.mod.check target
This switches to use `go.mod.check` target instead of `go.vendor.check` target.

Signed-off-by: moricho <ikeda.morito@gmail.com>
2020-03-17 16:48:43 +09:00
moricho 8907e220ea apis: regenerate files & modify k8sutil tests
This regenerate files with code-generator/deepcopy-gen and modify the related tests.

Signed-off-by: moricho <ikeda.morito@gmail.com>
2020-03-17 16:47:53 +09:00
moricho 2beb731a7e go: move to gomodules
This switches Rook to use `go mod` instead of `dep` for
dependencies management.

Signed-off-by: moricho <ikeda.morito@gmail.com>
2020-03-17 16:47:53 +09:00
Sébastien Han f44583b2b9 ceph: remove dead code
As part of the transition to support Ceph release **as of** Nautilus, we
left over that portion of code.
We don't need it anymore.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-16 14:01:24 +01:00