Commit Graph
43 Commits
Author SHA1 Message Date
parth-gr f007f2aca1 core: report node metrics using ceph telemetry
Add this reporting with the cephcluster reconcile,
Similar way we reported other telemetry's

Closes: https://github.com/rook/rook/issues/12344

Signed-off-by: parth-gr <paarora@redhat.com>
2023-11-16 19:13:22 +05:30
parth-gr 8e317ee074 ci: update golangci-lint version as it fails for some k8s version in 1.10
Closes: https://github.com/rook/rook/issues/11896

Signed-off-by: parth-gr <paarora@redhat.com>
2023-03-16 20:59:22 +05:30
Blaine Gardner b98efbb320 nfs: add support for running SSSD as a sidecar
Implement the SSSD sidecar design for NFS.

Because NFS documents are getting long, create an NFS
`Storage-Configuration` section with separated topics to keep the `NFS
Overview` document sane.

Adjust the SSSD design to allow any VolumeSource, not just ConfigMaps.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2022-08-11 14:59:23 -06:00
Travis Nielsen 5ef8d15659 build: format comments for go 1.19
The tool gofmt in go 1.19 requires certain formatting
in the comments section for better rendering.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2022-08-11 12:41:10 -06:00
Josh Soref 6e7b8767f3 core: fix spelling
* another
* are
* availability
* available
* bootstrap
* boundaries
* ceph
* certificate
* class
* codifies
* consuming
* corrupted
* createor
* csi
* deployments
* exceeded
* execute
* filesystem
* healthiness
* heuristics
* immediately
* insecure
* installed
* isolated
* maintained
* maximum
* minute
* monitor
* new
* nginx
* nonexistent
* not
* occurs
* omitempty
* operator
* orchestration
* persistentvolumes
* placement
* preexisting
* prometheus
* protecting
* provisioner
* purposes
* reconcile
* regex
* related
* requests
* returns
* rubbish
* running
* schedulable
* schedule
* serviceaccount
* simulating
* snapshots
* statement
* static
* tenants
* the
* unavailable
* volumeattachment
* waiting
* with
* wrapper
* zonegroup

Signed-off-by: Josh Soref <2119212+jsoref@users.noreply.github.com>
2022-07-07 18:10:47 -04:00
Alexander Trost 8686296e17 core: remove double imported packages
This removes double package imports. Example:
```
"github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
cephv1 "github.com/rook/rook/pkg/apis/ceph.rook.io/v1"
```
Only one is now being used as shown in go-staticcheck ST1019

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2022-04-25 13:51:45 +02:00
Sébastien Han b89730d895 ceph: refactor operator initialization sequence
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:

* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited

As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.

A second new controller for the operator's general config has been
created, it manages:

* the logging level
* the ceph CLI command timeout
* the discovery daemon

The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-09-17 16:57:12 +02:00
Blaine Gardner 1df4336a68 Merge pull request #7386 from BlaineEXE/update-osds-in-parallel
ceph: Update osds in parallel
2021-03-29 16:32:07 -06:00
Blaine Gardner 795124b7a8 ceph: update osds in parallel
Update OSDs in parallel per the design in
design/ceph/update-osds-in-parallel.md

The max number of OSDs updated in parallel is currently fixed at 20.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-03-29 10:55:28 -06:00
Blaine Gardner 71c41180be test: document or clean up codespell flags
Document codespell flags that are necessary. Remove codespell flags with
minor code changes if possible.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-03-26 15:00:49 -06:00
Satoru Takeuchi 4766954dfc ceph: suppress golanglint-ci complaints
Suppress the following complaints.

```
$ golangci-lint run -E gosec
pkg/operator/test/client.go:244:13: G404: Use of weak random number generator (math/rand instead of crypto/rand) (gosec)
        randIdx := rand.Intn(len(nodes.Items))
                   ^
tests/framework/installer/ceph_manifests_v1.5.go:44:19: G107: Potential HTTP request made with variable url (gosec)
        response, err := http.Get(url)
                         ^
pkg/daemon/ceph/client/pool.go:182:6: ineffectual assignment to stats (ineffassign)
        var stats = new(PoolStatistics)
            ^
pkg/operator/k8sutil/pod_test.go:34:2: ineffectual assignment to container (ineffassign)
        container, err := GetMatchingContainer([]v1.Container{}, expectedName)
        ^
```

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-03-12 15:34:27 +00:00
Blaine Gardner d5983e8d4b ceph: change osd on pvc unit test to use t.Log()
Change the TestOSDsOnPVC unit test to use t.Log()/t.Logf() to avoid
having a special infof() function that was unnecessary since our unit
tests run with `go test -v` where the `-v` flag interleaves t.Log()
output.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-02-25 11:52:00 -07:00
Blaine Gardner a95734160e ceph: create new osds on pvc before update existing
Add a new OSD provisioner status (reported by provisioner configmap)
that denotes that an OSD is preexisting and that OSD prepare does not
need to be run for the OSD. This only applies to OSDs on PVC currently
where existence of a deployment for the OSD indicates no further
provisioning needs to occur. New status is "preexisting".

Allow creating new OSDs on PVCs before updating existing ones by
allowing deferring OSDs during processing of OSD provisioning status
ConfigMaps. Deferral is identified by new "preexisting" status.

Add a unit test to ensure OSD on PVC provisioning occurs as expected
including deferring already-created OSDs to be updated after new PVCs
are created.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-02-12 17:38:35 -07:00
Arun Kumar Mohan 421f340c1c ceph: updating the dependencies for operator SDK v1.0.0
Updating the dependencies' versions to match with the newer Operator SDK
version v1.x

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:03:40 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
subhamkrai 25c116a4bd ceph: handle golangci-lint linter gosimple
this commit handles all the errors  check for
golangci-lint linter gosimple.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 15:10:27 +05:30
subhamkrai 465a0f0aec ceph: handling gosec error code g601
this commit handles all the gosec g601
error code (i.e Implicit memory aliasing
of items from a range statement).

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-28 11:34:49 +05:30
Travis Nielsen 631b13b906 ceph: refactor creds used by operator
The operator should only connect to ceph with a single set of creds.
In a converged cluster this will be the admin creds and in an external
cluster it will be lower-privileged creds. Independent clusters were
implemented with a separate set of creds. To simplify the code these
are now merged to a single set.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-16 15:54:42 -06:00
Sébastien Han bfa3b4daf4 ceph: do not use external node IP for mon endpoint
When the cluster is using HostNetworking, Rook was either picking up the
external or internal IP of the node. This resulted in mon endpoints have
public IP addresses. Those IP are not reachable from within the cluster
so OSD/CSI couldn't access the monitors from the configmap endpoint.

Also, exposing the cluster on a public network does not seem realistic,
so sticky with private/internal IP addresses is better.

Closes: https://github.com/rook/rook/issues/5495
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-04 14:12:45 +02:00
Nizamudeen 53883f68cf ceph: Handling Unhandled errors
This commit is to handle all those unhandled errors which raises the gosec warning.

Fixed G104: Unhandled Errors are handled now

Signed-off-by: Nizamudeen <nia@redhat.com>
2020-02-21 22:48:24 +05:30
Blaine Gardner 50e824ebad Ceph: Add unit tests for nfs-ganesha spec
Ganesha isn't a true "Ceph" daemon, so some of the sharable pod spec
testing functionality is teased out of the operator/ceph/test library
and a simple operator/test library is created. The ceph/test library is
updated to use the operator/test library when possible/appropriate.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-07-22 11:06:04 -06:00
Blaine Gardner a80082988d ceph: remove osds only when certain
Make the Ceph operator more cautious about when it decides to remove
nodes from the Rook-Ceph cluster which are acting as osd hosts.

When `useAllNodes` is set to `true` we assume that the user wants to
have the most hands-off experience. Node removals are allowed when a
node is delted from Kubernetes and when a node has its taints/affinities
modified by the user (but not by automatic k8s modification as much as
possible).

When `useAllnodes` is set to `false` the only time a node is removed is
if it is removed from the Ceph cluster definition.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-04-11 09:41:54 -06:00
Blaine Gardner f126887b52 tests: add T/F/Unk. status to node ready condition
Rook previously relied on the presence of `Ready` to determine if a node
is ready, but `Ready` can have different statuses. Correct this.

https://kubernetes.io/docs/concepts/architecture/nodes/#condition

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-04-09 15:40:37 -06:00
Blaine Gardner 35e2fbe4e0 optest: wait longer for deployment image
The Ceph upgrade test is the only one which uses the
`WaitForDeploymentImage` method, and it has to be configured to wait
longer after upgrade at this point. Since the wait time is still
hard-coded, this method is moved to the operator's test dir to make
it clear that the method is suitable only for tests currently.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-04-08 10:30:32 -06:00
Sébastien Han eb3c424d57 Ceph messengers 2 support
This commit introduces the necessary changes to support the new
messenger feature coming with Ceph Nautilus (currently in development).

What changes? Now the monitor listens on two port:

* old 6789 for messenger v1, which will help us support older client
(e,g: krbd)
* new 3300 for messengers v2, which brings new improvement in the
messaging layer. This new transport layer brings numerous advantages
such as encryption improvement, speed improvement, pluggable nature to
support different network stack than TCP and many more.

We still have one Service IP, however it has 2 ports, see:

```
NAME                      TYPE        CLUSTER-IP       EXTERNAL-IP   PORT(S)             AGE
rook-ceph-mon-a           ClusterIP   10.106.217.160   <none>        3300/TCP,6789/TCP   4h
rook-ceph-mon-b           ClusterIP   10.99.36.175     <none>        3300/TCP,6789/TCP   4h
rook-ceph-mon-c           ClusterIP   10.108.220.74    <none>        3300/TCP,6789/TCP   4h
````

The `ceph.conf` has changed and we don't force the port when using an IP
address (public addr etc). Ceph, depending on its version will naturally
start the monitors on their right port, 6789.

A new --ceph-version-name CLI argument has been added to the Rook binary
so that when the pod starts it passes the ceph version name and the
configuration of the ceph.conf, as well as daemon startup flags, happen
properly.

Given that the Rook Operator remembers the port of all the monitors it
deployed (through Pod definition), this change is not an issue and will
maintain backward compatibility.

Note that to test this you must build rook with dev container image,
which contains the dev Nautilus version. So you should do something
like:

`make -j4 BASEIMAGE='ceph/daemon-base:latest-master' IMAGES='ceph' build`

Resolves: #2525
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-02-22 00:09:37 +01:00
Blaine Gardner ec416cdd92 ceph mons: set up entirely in operator
Make the Rook config-init unnecessary for mons, and remove that init
container. Perform all mon configuration steps in the operator, and set
up the mon pods and k8s environment such that only Ceph containers are
needed for running mons.

This should help streamline changes to the mons, as there will be no
need to change the `daemon/mon` code or `cmd/rook/ceph` code with mon
changes in the future.

This work starts to lay the groundwork for supporting the
`design/ceph-config-updates.md` design.

Notable new bits:

Create a keyring secret store helper for storing dameon keyrings, and
use it to store the mon keyring. Mon pods mount the keyring into a
k8s secret-backed volume.

Create a configmap store for the Ceph config file which can be mounted
into pods/containers directly to /etc/ceph/ceph.conf. Also store
individual mon_host and mon_initial_members values which can be mapped
into pods as environment variables and used in Ceph commandline flags,
enabling the mon pods to have the most up-to-date information about the
mon cluster when restarting and without need for operator intervention.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-02-05 06:19:59 -07:00
travisn c2f354680e mons: use the default port of 6789 instead of 6790
Signed-off-by: travisn <tnielsen@redhat.com>
2019-01-17 14:37:58 -07:00
Blaine Gardner ce803fbf9d Ceph mds/file: Set up config in init container
Progress toward issue #2003.
Includes design from design doc PR #1578

Use init containers to create configuration for Ceph mgrs. There is only
one init container in this design. The init container calls the Rook
binary to create Ceph config files which are then shared with the mds
daemon main container.

Once this init is run, the main mds daemon is run. Leaving room to use
the Ceph-versioned image in the future, call `ceph-mds --foreground ...`
to run the Ceph mds.

The refactor to using an init container also necessitated refactoring
the mdses replicaset implementation to a deployment-per-pod
implementation due to a chicken-egg problem. With a single container (in
the before times) the Rook binary was able to call the ceph-mds daemon
with an id generated from the pod name. Since the pod name is not known
before runtime, and the id is one of the few params that must be
specified to Ceph daemons on run, it is necessary to know the id
beforehand; thus the move to a deployment architecture following the
likes of the mon and mgr daemons.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-10-05 19:01:12 -06:00
Blaine Gardner ca71b7a311 Ceph mgr: Set up config in init container
Progress toward issue #2003.
Includes design from design doc PR #1578

Use init containers to create configuration for Ceph mgrs. There is only
1 init container in this design:
 1. Using the Rook image, call the Rook binary to create Ceph config
files shared with the mgr daemgr container.

Once this init is run, the main mgr daemgr is run. Leaving room to use
the Ceph-versioned image in the future, call `ceph-mgr --foreground ...`
to run the Ceph mgr.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-09-21 13:11:08 -06:00
Blaine Gardner 7a5fe8498c Ceph op: Add test package, container tests
Add a 'test' package to the Ceph operator and define a test to verify
that containers produced by Rook-Ceph match what is expected. Because
this is for 'containers' and not strictly for 'ceph containers', this
could be moved a level up to the operator test package; however, if
other backends wish to use the container tests, they will likely need to
make modifications, and there is concern that this might make the Ceph
tests brittle.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-09-21 09:58:32 -06:00
Blaine Gardner 78c2cec1ad Add k8s volume/mount tests to operator test pkg
Add functions for helping test Kubernetes volumes and volume mounts.

- Create functions for testing the existence of volumes/mounts by
name in a list of vols/mounts without needing to know the index of
the vol/mount in the list.

- Create functions for printing vols/mounts in a human-readable format
so that tests may output more useful errors.

- Create a test definition for ensuring that all the volumes a pod
provides and all the mounts from each pod container match. For each
volume, there must be at least one container which mounts the volume,
and for each mount, there must be a source volume.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-09-21 09:58:31 -06:00
Blaine Gardner 3731f06176 Ceph: Refactor mon config/keyring to cephconfig
Create a cephconfig module in Ceph's daemon pkg source, and refactor the
config and keyring generation that exists in the mon package into the
new cephconfig package. The config/keyring generation code is used by
most all daemons and not just mon, so a new package is a more
appropriate place for this.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2018-09-05 10:17:11 -06:00
travisn 5f01178e83 mon: set names based on chars instead of integers
Signed-off-by: travisn <tnielsen@redhat.com>
2018-08-30 22:18:53 -06:00
Alexander Trost 86568cafd3 operator/ceph/mon: Cleaned up tests
To make it more applicable to what is actually running when using the
operator the used test mon names have been replaced by the actual format
used in normal operations: `rook-ceph-mon[0-9]` instead of `mon[0-9]`.

This also removes some duplicated test code and used the already
available function for it.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2018-06-04 23:01:29 +02:00
Travis Nielsen 8725170077 operator: mon affinity to nodes based on node hostname 2018-01-17 14:47:22 -08:00
Travis Nielsen 0298d1fb29 completed conversion for k8s 1.8 client conversion
move core rados packages under cluster folder

move daemons under daemon pkg

rename flex crd package to attachment

pin dep to k8s 1.8.2
2017-12-11 14:35:07 -08:00
Travis Nielsen 4e6774c063 remove legacy standalone code 2017-10-17 12:23:57 -07:00
Alexander Trost 4f4f727bcd Added hostNetwork option to cluster spec 2017-09-18 19:57:20 +02:00
Travis Nielsen f3640887f3 golint cleanup for the operator 2017-07-24 17:24:24 -07:00
Steve Leon 668ed5d25b Incorporating new Kubernetes API changes
- This is needed to make it compatible with k8s 1.8
2017-07-13 11:08:07 -07:00
Travis Nielsen c122562a7d Rook refactor to call external ceph tools instead of embedded ceph
Remove all cgo from the code base, and start the transformation to calling external ceph tools to generate config and start daemons.
2017-06-20 20:57:45 -07:00
Travis Nielsen 4710b63e65 mons are a replicaset instead of started directly as a pod 2017-06-05 15:14:10 -07:00
Travis Nielsen e0e829ed25 Rook operator unit tests 2017-03-09 15:24:33 -08:00