Commit Graph
460 Commits
Author SHA1 Message Date
Jiffin Tony Thottan da61c9a83e ceph: vault kms configuration for ceph object store
The first patch to configure vault for ceph object store. If the `security.kms` configured in
`clusterSpec` CRD, RGW will be configured with vault kms settings to handle SSE request from s3 clients.

Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
2021-01-22 10:44:03 +05:30
Sébastien Han 362637ec3d ceph: do not explicitly turn on the balancer
As of Pacific the balancer is now on by default in upmap mode.
In earlier versions, the balancer was included in the `always_on_modules` list, but needed to be
turned on explicitly using the ``ceph balancer on command.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-01-08 17:20:47 +01:00
shenjiatong 8d3d460e70 ceph: refactor and remove duplicate codes
Signed-off-by: shenjiatong <yshxxsjt715@gmail.com>
2020-12-18 18:01:30 +08:00
shenjiatong 04aabb3896 ceph: analyze result for new typed return from lvm batch
Signed-off-by: shenjiatong <yshxxsjt715@gmail.com>
2020-12-18 11:11:03 +08:00
ushen 373b952988 ceph: allow ceph 14.2.15 for lvm batch
14.2.15 lvm batch command prepare report changes
output format. This commit skips md check if ceph version
is greater than 14.2.13.

Signed-off-by: shenjiatong <yshxxsjt715@gmail.com>
2020-12-18 11:11:02 +08:00
Travis Nielsen 45e8ca7173 Merge pull request #6838 from synarete/ceph-crush-rule-with-device-class
ceph: Two step crush rule with device class
2020-12-16 13:08:52 -07:00
Travis Nielsen 64fc740d8a ceph: revert new node is not provisioned when useallnodes is true
This reverts commit a6ee5657ae.
The operator reconcile of the cephcluster was staying in an endless
loop in clusters that did not have OSDs configured on all nodes.
Now we revert this change for the patch release and will follow
up with a more complete fix in https://github.com/rook/rook/pull/6813.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-16 11:36:31 -07:00
Shachar Sharon 39a806f747 test: ceph add-rule via crush-map explicit decompile/compile
Test the code-flow when creating crush-rule via explicit
decompile/compile crushmap, when building two-steps crush rule. Relevant
for the case where 'PoolSpec.Replicated.ReplicasPerFailureDomain' is non
zeoro.

This test requires ceph's 'crushtool' to be installed on local machine.
Otherwise, it is ignored.

Signed-off-by: Shachar Sharon <ssharon@redhat.com>
2020-12-16 17:43:49 +02:00
Shachar Sharon 53461dd2ef ceph: pass 'device-class' when creating explicit crush-rule
When 'pool.Replicated.ReplicasPerFailureDomain' is non zero, the code
goes via route of adding explicit crush-rule from template, but ignores
the 'pool.DeviceClass' if it is set. Added option to template to have
(optional) 'class hdd|ssd'

Signed-off-by: Shachar Sharon <ssharon@redhat.com>
2020-12-16 17:42:27 +02:00
Travis Nielsenandshenjiatong a0a27bc426 ceph: apply deviceClass properly to new osds
The deviceClass property was being ignored when creating the
non-pvc OSDs. Now the deviceClass will be specified as a property
for individual devices, all devices on a node, or all OSDs in the
cluster, depending on the level where the config is applied in the
cluster CR.

Co-authored-by: shenjiatong <yshxxsjt715@gmail.com>
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-11 10:16:00 -07:00
Travis Nielsen d0ba899341 ceph: check stretch cluster is ready to configure arbiter
The arbiter can only be configured with the stretch cluster if the
CRUSH map is balanced and there are two zones in the CRUSH map.
After the OSDs are configured, we wait for all the OSD pods to be
running and that the CRUSH map is balanced. If it takes more than
two minutes, we fail the reconcile and try again. This is only done
the first time the stretch cluster is configured. In future reconciles
we first check if the stretch cluster is already enabled before
enabling it again.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-12-09 16:37:24 -07:00
Travis Nielsen 6fb3c522bb Merge pull request #6768 from shaas/master
ceph: fix #6214 New node is not provisioned when useAllNodes: true
2020-12-04 08:57:49 -07:00
Stefan Haas a6ee5657ae ceph: fix #6214 New node is not provisioned when useAllNodes: true
Signed-off-by: Stefan Haas <shaas@suse.com>
2020-12-04 15:05:58 +01:00
Sébastien Han aca5a8cc53 ceph: use debug message for snap schedule
We don't need to print this every 60s in the operator log.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-26 09:12:11 +01:00
Sébastien Han 5d9612c2ae ceph: fix metadata device passed by-id
The code was assuming that devices were passed by the user as
"/dev/sda", this is bad! We all know people should be using paths like
/dev/disk/by-id so we must support them.

Closes: https://github.com/rook/rook/issues/6685
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-25 15:00:47 +01:00
Arun Kumar Mohan ded16f779d ceph: changes for 'sigs.k8s.io/sig-storage-lib-external-provisioner/v6'
Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:03 +05:30
Arun Kumar Mohan 65d16bfc94 ceph: manual changes needed for kubernetes api updates
Fetched latest lib-bucket-provisioner changes as well.

Signed-off-by: Arun Kumar Mohan <amohan@redhat.com>
2020-11-18 21:14:01 +05:30
Sébastien Han afc7ecff31 ceph: add snapshot scheduling for mirrored pools
Now, we can schedule snapshots on pools from the CephBlockPool CR when
the pool is mirrored.
It can be enabled like this:

```
mirroring:
  enabled: true
  mode: pool
  snapshotSchedules:
    - interval: 24h # daily snapshots
      startTime: 14:00:00-05:00
```

Multiple schedules are supported since snapshotSchedules is a list.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-18 09:50:09 +01:00
Lalit Maganti 1d1a9f111a ceph: associate newly created pools to existing filesystem
When new pools are added to an already existing filesystem, we would
create them but not actually associate them to the filesystem. Ensure
that we do this.

Also, while we're here, we refactor the CreateFilesystem test: before,
we were not properly testing creating the filesystem from scratch
because we would return valid JSON for *every* ceph command. Change this
by making the first run return an error so that we get some better
coverage.

After this, we also add a check for the correctness of adding new pools
to the CreateFilesystem unittest.

Fixes #5876

Signed-off-by: Lalit Maganti <lalitm@google.com>
2020-11-11 22:16:37 +00:00
Sébastien Han e3d032dbc1 Merge pull request #6489 from travisn/stretch-mons
ceph: Stretch cluster configuration
2020-11-06 15:16:39 +01:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
Sébastien Han 8631eef1f7 ceph: correctly populate variable
The wrong variable was referenced in the print.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-05 18:06:47 +01:00
Sébastien Han bc5c8d32f1 ceph: fix crush-class osd on non-PVC
We need to use storeConfig instead device config as the later never
seems to be populated.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-05 18:06:47 +01:00
Sébastien Han 6b211eacba ceph: add unit test for osd batch mode
We now have much better coverage on the c-v on non-pvc scenario.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-05 18:06:11 +01:00
Travis Nielsen e72b69f8e0 ceph: validate pools for stretch clusters
Pools in stretch clusters will all use the same crush rule that
is generated by the operator when configuring the cluster for
stretch mode, with replica 4 and two replicas per failure domain.
Pools cannot create new crush rules, so we require that replica be
4 when the pools is created, EC is not allowed, and no new
crush rule will be created.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-05 09:23:57 -07:00
Travis Nielsen b91f4211c9 ceph: configure a stretched cluster
In clusters where only two datacenters (or similar failure domains)
are available, a different mon and osd approach is needed to deal
with the network partitions or some other reason for one of the failure
domains going down. The Ceph stretched cluster makes the mons aware
of the failure domains by configuring one as the arbiter in a third
zone, while keeping two replicas of the data in each of the data
zoens.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-05 09:17:52 -07:00
Travis Nielsen 37a56f39a6 ceph: remove the osd pvc when purging the osd
The osd purge job was not removing the osd prepare job or the
osd pvc due to an invalid label query for the pvc. Now the
prepare job and pvc will be removed as expected from the job.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-11-03 08:32:37 -07:00
Sébastien Han b7e0e206d3 ceph: rework unit test
Golang maps order is not always guaranteed to be identical, which makes
the table driven test to fail. Because of the nature of the maps, it is
hard to predict the exact position of an element in the map. This
makes the table driven test flappy.
Reworked the unit test using a contains statement which achieves the
same validation.

Closes: https://github.com/rook/rook/issues/6522
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-11-02 12:14:09 +01:00
Sébastien Han ea1d71cbfb ceph: add vault kms support for osd encryption
When the Ceph cluster runs on PVC and the OSDs are encrypted we can
store LUKS's Key Encryption Key inside a Key Management System. Today,
Rook only supports HashiCorp Vault: https://www.vaultproject.io/

The CephCluster has now a new "security" field which will plug onto the
KMS. Here is an example:

security:
  kms:
    tokenSecretName: <name of the secret containing a Vault token, used
    to authenticate>
    connectionDetails: < a map of strings containing connection
    information>

Refer to the ceph-cluster-crd documentation to lear more.

Closes: https://github.com/rook/rook/issues/6105
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-30 16:16:33 +01:00
Jonas Schäfer 7ace6ac255 ceph: make root= CRUSH label value configurable
The CRUSH map does not allow duplicate values (nor keys) for the
labels. That means that if the currently hardcoded label value
`default` occurs anywhere else in the tree (for example, because
some of the topology labels provided by the cloud provider use it),
the cluster cannot sucessfully spawn.

By giving the user control over the label value used for the root
CRUSH map label, it is possible to adapt the cluster to this
situation.

To be able to actually use the `default` value, the root=default
node which is created by default by ceph itself has to be removed;
since that node is accompanied by a replicated crush rule, we have
to remove that rule, too.

The integration tests are modified to sometimes use the custom
crushRoot (based on an arbitrary criterium) to get coverage across
the various suites. Separate tests could be added, but were not
deemed necessary at this point.

Fixes #4993.

Signed-off-by: Jonas Schäfer <jonas.schaefer@cloudandheat.com>
2020-10-29 09:22:41 +01:00
subhamkrai cb0ca66a6b ci: enable gosec linter in golangci-lint
golangci-lint linter gosec showing more errors
than gosec gh. This commit resolve
new errors. And, removing
nosec comments from autogenerated files.

Signed-off-by: subhamkrai <srai@redhat.com>
2020-10-19 11:50:06 +05:30
Sébastien Han f0b2c6eacf Merge pull request #6184 from cybozu-go/ceph-raw-mode-on-lv-backed-pvc
ceph: raw mode OSD on LV-backed PVC
2020-10-15 09:46:56 +02:00
Daichi Sakaue b9f29349ed ceph: use raw mode instead of lvm mode even if lv-backed PVC
It can simplify OSD management in many ways and can avoid many problems
that come from LVM.

Closes: https://github.com/rook/rook/issues/5939
Signed-off-by: Yuji Ito <llamerada.jp@gmail.com>
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2020-10-14 01:46:25 +00:00
Sébastien Han b70b098405 ceph: add support for stretched cluster crush rule
The pool spec has now a new property called "replicasPerFailureDomain"
which essentially represents the number of replicas to store in each
failure domain.
Assuming the failure domain is a datacenter (if the cluster is
stretched) then you will have 2 replicas per datacenter where each
replica ends up on a different host. This gives you a total of 4
replicas and for this, the "size" must be set to 4.

Closes: https://github.com/rook/rook/issues/5591
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-13 16:33:34 +02:00
Sébastien Han b623133fb6 ceph: add ceph-volume log to non-pvc scenario
When deploying on non-PVC, we still need to gather c-v logs on failures.
Now if the bootstrap fails, the c-v log will be printed.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-12 14:10:17 +02:00
subhamkrai 0fddfcf307 ceph: handle golangci-lint linter errcheck error
this commit handle golangci-lint linter errcheck.

`errcheck` - Errcheck is a program for checking for
unchecked errors in go programs. These unchecked errors
can be critical bugs in some cases

To see only staticcheck linter output
`golangci-lint run --disable-all -E errcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-30 22:24:34 +05:30
subhamkrai de8dbbcdcc ceph: handle golangci-lint linter staticcheck error
this commit handle golangci-lint linter staticcheck error.

`staticcheck` - Staticcheck is a go vet on steroids,
applying a ton of static analysis checks.

To see only `staticcheck` linter output

`golangci-lint run --disable-all -E staticcheck`

Signed-off-by: subhamkrai <srai@redhat.com>
2020-09-24 15:04:55 +05:30
subhamkrai 829778f251 ceph: handle golangci-lint linter ineffassign
this commit will enable one more linter ineffassign
in golangci-lint.

This linter throws an error when variable is assigned and never used.

`golangci-lint run --disable-all -E ineffassign` is used detects ineffassign
errors only.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-21 14:31:18 +05:30
subhamkrai 1e9bb8e6e2 ceph: handle golangci-lint linter deadcode
this commit will enable one more linter `deadcode`
in golangci-lint .

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-18 16:53:49 +05:30
Sébastien Han d118bd38a6 Merge pull request #6278 from subhamkrai/golanci-lint-unused
ceph: handle golangci-lint linter unused
2020-09-18 12:44:37 +02:00
Sébastien Han 204ffd4065 Merge pull request #6250 from leseb/mirroring-config
ceph: add rbd-mirror configuration
2020-09-18 09:16:08 +02:00
subhamkrai f9fafe62d4 ceph: handle golangci-lint linter unused
this commit will enable one more linter
in golangci-lint.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 22:53:24 +05:30
Sébastien Han 451622a955 ceph: add rbd-mirror configuration
Rook is now capable of configuring mirroring between sites. The
implementation works at different levels:

* CephBlockPool: which introduces a new `mirroring` configuration as well
as `statusCheck`. When turned on, Rook will enable mirroring on the
pool. It will also create a bootstrap peer token and store it in a
Kubernetes Secret. The name of that Secret can be found in the Status
field of the CephBlockPool CRD. This token can be fetched and used by
other clusters to configure the site as a peer. Mirroring can be
configured either at the pool or the image level.

* CephRBDMirror: which introduces a new `peers` configuration allowing
Rook to connect to peers by passing a Secret name. The administrator will
create a Kubernetes Secret with 2 keys: 'token' for the bootstrap peer
token and 'pool' for the name of pool. Once detected the rbd-mirror
controller will go ahead and import the peer configuration.

Pool mirroring status example:

```
status:
  info:
    rbdMirrorBootstrapPeerSecretName: pool-peer-token-test
  mirroringInfo:
    lastChanged: "2020-09-17T14:47:27Z"
    lastChecked: "2020-09-17T14:48:27Z"
    summary:
      summary:
        mode: image
        peers:
        - client_name: client.rbd-mirror-peer
          direction: rx-tx
          mirror_uuid: ""
          site_name: rhcs
          uuid: c50522a4-28a4-4bd3-ba68-e11780308882
        site_name: 91eae0dd-06b1-4d2c-91f3-1311c9df382b-rook-ceph
  mirroringStatus:
    lastChecked: "2020-09-17T14:48:27Z"
    summary:
      summary:
        daemon_health: OK
        health: OK
        image_health: OK
        states:
          replaying: 1
```

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-09-17 19:05:04 +02:00
subhamkrai 25c116a4bd ceph: handle golangci-lint linter gosimple
this commit handles all the errors  check for
golangci-lint linter gosimple.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-09-17 15:10:27 +05:30
n.fraison de7ca01930 ceph: merge rook-config-override to the default global config file
Some parameter like bluestore_min_alloc_size, bluestore_min_alloc_size_hdd need to be set during the bootstrap of the
ceph cluster in order to have bluestore created accordingly.
To do so rook users rely on the rook-config-override configmap.
Sadly the rook-ceph-osd-prepare job does not take in account this configmap to generate its ceph config.
Changing the behaviour to merge the config from rook-config-override configmap to the default ceph config

Signed-off-by: n.fraison <n.fraison@criteo.com>
2020-09-16 15:08:50 +02:00
Servesha Dudhgaonkar 2cf0518545 ceph: update remove.go
Call logger.Infof() instead of Debugf() in remove.go file.

Signed-off-by: Servesha Dudhgaonkar <sdudhgao@redhat.com>
2020-09-09 20:06:27 +05:30
Servesha Dudhgaonkar 0c323439a6 ceph: purge a down osd with a job
If an OSD is down and needs to be removed from the cluster,
a job can be run that will remove the OSD deployment
and purge the OSD from ceph. If the OSD is still up,
the purge will be rejected.

To purge multiple OSDs at the same time, the OSD IDs can be
specified as a comma-separated list.

Signed-off-by: Servesha Dudhgaonkar <sdudhgao@redhat.com>
2020-09-01 09:36:20 +02:00
Sébastien Han f2313d6db0 ceph: do not exec on rootfs just lookup
Trying to execute a binary from the host within the container can result
in mismatched libraries. The host version is compiled with paths of
libraries on the host where the libs inside the container can differ.
Because the path is known and we only need to make sure whether the bin
exists or not then simply doing a lookup when we fallback on /rootfs is
sufficient.

Closes: https://github.com/rook/rook/issues/6078
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-08-27 15:16:46 +02:00
Sébastien Han aa5362d48d Merge pull request #6120 from leseb/fix-5996
ceph: do not use realPath for osd on pvc
2020-08-19 17:33:46 +02:00
Sébastien Han 944d8cb209 ceph: do not use realPath for osd on pvc
It is fine to use /mnt/set1-data-0-k47mc, it looks like I thought there
was an issue with udev but there is none...

Closes: https://github.com/rook/rook/issues/5996
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-08-19 14:39:55 +02:00