Commit Graph
2956 Commits
Author SHA1 Message Date
andrew webber (personal) 2eaaf26792 Container Linux - Daemonsets missing required field "selector"
Signed-off-by: andrew webber (personal) <andrewvwebber@googlemail.com>
2019-06-24 23:36:40 +02:00
Travis Nielsen 818ad9c7ca Merge pull request #2939 from travisn/delay-system-daemons
Delay starting the Rook system daemons until a CephCluster CR is created
2019-06-18 16:52:12 -06:00
Travis Nielsen 1d35f26302 Merge pull request #3313 from leseb/osd-sdn
ceph: osd: fix startup on sdn
2019-06-18 16:13:52 -06:00
Alexander Trost a0df634dad remove unused md file (#3301)
remove unused md file
2019-06-18 17:59:22 +02:00
Alexander Trost aa92e9a653 ceph: add psp.yaml example (#3317)
ceph: add psp.yaml example
2019-06-18 17:48:29 +02:00
Jared Watts 4d53b88cf1 Merge pull request #3254 from jbw976/governance-code-approvers
governance: new change approval process to include approvers/reviewers
2019-06-18 08:18:13 -07:00
Sébastien Han a74f806153 ceph: add psp.yaml example
The current doc was outdated and hardcoded, so removing the example from
the doc and add a proper psp.yaml file that people can use and
contribute too.

Closes: https://github.com/rook/rook/issues/3309
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-18 16:58:28 +02:00
Travis Nielsen 9df4b24197 Merge pull request #3117 from rohan47/osd_marked_out
Clean up the OSD after the OSD is marked "out"
2019-06-18 08:13:35 -06:00
Travis Nielsen 4fd227a3cc Merge pull request #3319 from ShyamsundarR/update-3312-imageversion
Update openshift with Ceph-CSI image versions as in Kubernetes case
2019-06-18 08:00:47 -06:00
ShyamsundarR 61b9f9a9f0 Update openshift with Ceph-CSI to image versions as in Kubernetes case
PR #3217 changed the pod manifest to drop some parameters to the
ceph-csi pods. This also resulted in a change to the operator with CSI
yaml for non-openshift case, but failed to update similar yaml's for
the openshift case.

This commit rectifies this problem.

Updates: #3312
Signed-off-by: ShyamsundarR <srangana@redhat.com>
2019-06-18 07:07:00 -04:00
Sébastien Han b2f2ff6b9d ceph: osd: fix startup on sdn
This commit adds a new flag to the osd startup so that on msgr2 (default
on Nautilus and above) the osd is able to find the IP address in the
container to bind to.

This requires this Ceph patch https://github.com/ceph/ceph/pull/28589
and is already present in Octopus.

Closes: https://github.com/rook/rook/issues/3140
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-18 10:19:19 +02:00
travisn c6c4a9b42a ceph: delay starting the system daemons until a cluster is created
When the operator first starts, the only operation needed
is to watch for new cephcluster crds to be created and
start the discovery to find available devices. The flexvolume
agent, the csi driver, and the volume provisioning can all be delayed
starting until the first cluster is created.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-06-17 15:44:33 -06:00
Jared Watts d9f11a1b25 governance: new change approval process to include approvers/reviewers
Signed-off-by: Jared Watts <jbw976@gmail.com>
2019-06-16 20:59:32 -07:00
Travis Nielsen 8da2221296 Merge pull request #3283 from leseb/rgw-refactor
ceph: refactor rgw bootstrap
2019-06-14 15:17:58 -06:00
Travis Nielsen 488fe64815 Merge pull request #3281 from phantooom/master
Fix onDelete func panic
2019-06-14 11:59:35 -06:00
Travis Nielsen 831106cb3e Merge pull request #3271 from phlogistonjohn/jjm-ceph-csi-rdb-mon-cfg-map
ceph csi: Create and maintain mon config to be consumed by csi
2019-06-14 11:53:41 -06:00
rohan47 a420db69fe osd: Clean the osds that are out and safe-to-destroy
- When an osd is marked out, and it is safe to be destroyed then delete the osds deployment.
- Removed osdGracePeriod
- Updated unit tests to test osd marked out action

Signed-off-by: rohan47 <rohgupta@redhat.com>
2019-06-14 22:50:33 +05:30
Sébastien Han 93b2448619 ceph: refactor rgw bootstrap
This commit does multiple things:

* remove support for AllNodes where we would deploy one rgw per node on
all the nodes.
* a transition path is implemented in the code so that if someone has an
existing deployment, daemonsets will be removed and replaced by an
deployments.
* when using "instances", each rgw deployed has its own key which makes
Ceph reporting the exact number of rgw running, see:

```
[root@rook-ceph-operator-775cf575c5-bh4sr /]# ceph -s
  cluster:
    id:     611fcf39-0669-4864-9a12-debb35c0397a
    health: HEALTH_OK

  services:
    mon: 3 daemons, quorum a,b,c (age 12h)
    mgr: a(active, since 12h)
    osd: 3 osds: 3 up (since 12h), 3 in (since 12h)
    rgw: 3 daemons active (my.store.a, my.store.b, my.store.c)

  data:
    pools:   6 pools, 600 pgs
    objects: 235 objects, 3.8 KiB
    usage:   3.0 GiB used, 84 GiB / 87 GiB avail
    pgs:     600 active+clean
```

Closes: https://github.com/rook/rook/issues/2474, https://github.com/rook/rook/issues/2957 and https://github.com/rook/rook/issues/3245
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-14 17:55:15 +02:00
John Mulligan 214d08718e doc: Update csi document to match new storageclass field
Update the csi document to assist the user in updating the new
`clusterID` field that is used to map to mons in the configmap
rook maintains.

Signed-off-by: John Mulligan <jmulligan@redhat.com>
2019-06-14 10:45:50 -04:00
John Mulligan 8c5a1c502e ceph csi: update storage class example with clusterID field
New versions of ceph csi rbd expect a clusterID that will be used
to index into the config map and determine what mons to use.
Also, remove mons from config example.

Signed-off-by: John Mulligan <jmulligan@redhat.com>
2019-06-14 10:45:50 -04:00
John Mulligan 431f504cd4 ceph: create and maintain a config map for ceph csi to consume
Create and maintain a config map that meets the requirements of
the ceph csi such that Rook can maintain the contents of config
map with up-to-date mon information to be used later by csi.

Signed-off-by: John Mulligan <jmulligan@redhat.com>
2019-06-14 10:45:41 -04:00
John Mulligan d8ba51084b ceph csi: update config templates to match "canary" csi version
This is a temporary change that fixes the templates so that they match
the so-called "canary" tag in csi. This version of the csi rbd driver
supports a external mon configuration (in a config map).

Signed-off-by: John Mulligan <jmulligan@redhat.com>
2019-06-14 10:43:53 -04:00
John Mulligan 700b42a475 ceph csi: update the operator-with-csi.yaml to use canary csi images
Temporary change to support the testing and development of new
integration between ceph csi and rook. This "canary" tag points
at new versions of csi that support taking mon config from a
config map.

Signed-off-by: John Mulligan <jmulligan@redhat.com>
2019-06-14 10:43:53 -04:00
Sébastien Han f6f1aa772f rgw: remove legacy code
This code can be removed since 1.0 shipped.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-14 16:28:17 +02:00
Travis Nielsen d44eca72e8 Merge pull request #3296 from leseb/nautilus-osds
ceph: upgrade apply osd nautilus flag
2019-06-14 07:28:32 -06:00
xiaorui.zou c4aef1a073 remove unused md file
Signed-off-by: xiaorui.zou <xiaorui.zou@gmail.com>
2019-06-14 13:42:47 +08:00
Sébastien Han cba9a359a0 ceph: upgrade apply osd nautilus flag
When OSDs are running on Nautilus we always disable old osd features and
aplpy the onces for Nautilus as described in the upgrade doc.
During an upgrade or the next time an orchestration will be called the
command will be applied. The command is idempotent so we can run it each
time.
This can be backported for 1.0.3

Closes: https://github.com/rook/rook/issues/2960
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-13 22:15:15 +02:00
Travis Nielsen 7d3878378d Merge pull request #3259 from noobaa/guy-rook-noobaa-design
Rook-NooBaa Design Doc
2019-06-12 21:45:03 -06:00
xiaorui.zou b3cd67f871 Fix onDelete func panic
onDelete obj in some case will return DeletedFinalStateUnknown type, we need use assert.

Signed-off-by: xiaorui.zou <xiaorui.zou@gmail.com>
2019-06-13 09:31:53 +08:00
Travis Nielsen 0f1c919b83 Merge pull request #3255 from rhcs-dashboard/ceph-dashboard-enable-object-gateway
ceph: updated doc. for enabling dashboard object gateway mgmt.
2019-06-12 15:45:02 -06:00
Travis Nielsen fc25c6d1ca Merge pull request #2675 from ashishranjan738/path
ceph: enhance server to search for rookflex
2019-06-12 15:44:23 -06:00
Ashish Ranjan 5a19ab9545 ceph: enhance server to search for rookflex
Signed-off-by: Ashish Ranjan <ashishranjan738@gmail.com>

This commit enables server to search for `rookflex` binary instead of assuming it to be present in `/usr/local/bin/`.

Fixes: https://github.com/rook/rook/issues/2486
2019-06-12 23:36:09 +05:30
Guy MargalitandTravis Nielsen fe5d94e370 Rook-NooBaa Design Doc
Signed-off-by: Guy Margalit <guymguym@gmail.com>
Co-Authored-By: Travis Nielsen <tnielsen@redhat.com>
2019-06-12 01:44:04 +03:00
Travis Nielsen 14e68dd5ce Merge pull request #3275 from iMartyn/patch-1
[skip-ci] Include AKS in the oddity list
2019-06-10 10:10:54 -07:00
Travis Nielsen 7e093cb18a Merge pull request #3280 from leseb/change-rgw-backend
rgw: change default frontend on nautilus
2019-06-10 08:27:57 -07:00
Sébastien Han 0317de9096 rgw: change default frontend on nautilus
As per: ceph/ceph#26599, Beast is now the
default fronted for rados gateway.
Newly created cluster as of Nautilus will use it by default.

Re-added version of 03587352d5
Resolves: #2707
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-07 23:10:11 +02:00
Travis Nielsen f2c49ea034 Merge pull request #3165 from d-luu/resource_comparer
ceph: added comparer for resource quantity when checking cluster changes
2019-06-07 11:15:33 -07:00
Alexander Trost aff2874f17 Merge pull request #3276 from leseb/fix-doc
doc: fix patch command
2019-06-07 19:49:56 +02:00
Travis Nielsen 40ea3e65a2 Merge pull request #3274 from dyusupov/master
Enable proper usage of metadataOnly property
2019-06-07 10:45:03 -07:00
Travis Nielsen d73ff85b51 Merge pull request #3257 from leseb/retry-version-detect
ceph: retry on detecting ceph version
2019-06-07 09:44:26 -07:00
Sébastien Han e0584f344d doc: fix patch command
The command had the namespace twice, remove the typo and used a working
command.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-07 18:16:10 +02:00
Sébastien Han 6438d9907e ceph: retry on detecting ceph version
Sometime the kube engine needs a bit of time to return the logs of a
given job and fails to read the stream.
Retrying up to detect the Ceph version seems reasonnable to
overcome this issue.

Fixes: https://github.com/rook/rook/issues/3227
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-07 18:13:38 +02:00
Travis Nielsen 6d1ccadd4b Merge pull request #3256 from noahdesu/device-hotplug-update
discover: handle false-positives observed by users
2019-06-07 08:55:00 -07:00
Travis Nielsen c0276802f5 Merge pull request #2564 from travisn/toplevel-fsgroup
Set fsgroup on the top level of the mount
2019-06-07 08:53:59 -07:00
Noah Watkins 3966f163ed discover: handle false-positives observed by users
the only exception to a naive device list comparison had been to ignore
drive UUID information which was unreliable when a device wasn't
formatted / partitioned. however various users have reported different
type of false positives that resulted in orchestration being run
continuously due to the wrong observation that devices were changing.

this patch fixes the cases we have observed and attempts to be slightly
more conservative in the calculation.

1. the devlinks is ignored. when a device is setup for lvm, for example,
the devlinks will be updated with different paths that point to the
device in addition to its standard paths addressable by pci address.

2. in the lvm case, the "model" field and "filesystem" field may also
change.

3. we ignore devices with devlinks that contain "usb" to avoid issues
when using usb drives.

4. be smart about detecting device availability. if a device transitions
from a non-empty (or has-partitions) state to an empty (or unpartitioned)
state then orchestration is triggered. this like observing that a device
is now available (e.g. in the allDevices case). however, when a device
transistions from empty to non-empty, then this is ignored as while it
is a change, it's generally a change associated with the new consumption
of the device.

fixes: #3059
fixes: #3185
fixes: #3131

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-06-07 08:22:08 -07:00
Travis Nielsen 8cf5d540a4 Merge pull request #3258 from leseb/doc-resource-limit
Doc resource limit
2019-06-07 07:52:15 -07:00
Martyn Ranyard 397904ee14 [patch-1] Include AKS in the oddity list
Signed-off-by: Martyn Ranyard <m@rtyn.berlin>
2019-06-07 15:56:38 +02:00
Sébastien Han 62b69d08e5 ceph: add missing documentation for resource limits
This commit documents the minimum amount of memory accepted by Rook to
 run Ceph pods properly.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-07 15:06:16 +02:00
Sébastien Han d3c8d5612b ceph: add resource limit check for rbdmirror
The memory check was missing and will be trigger if resources limit are
configured for the rbdmirror pod.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-07 15:06:16 +02:00
Sébastien Han 24abbddf14 ceph: fix pod memory check
This commit fixes the second test case where limit and request are
either identical or different but still we use limit as a value.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-07 15:06:16 +02:00