Commit Graph
312 Commits
Author SHA1 Message Date
Sébastien Han cba9a359a0 ceph: upgrade apply osd nautilus flag
When OSDs are running on Nautilus we always disable old osd features and
aplpy the onces for Nautilus as described in the upgrade doc.
During an upgrade or the next time an orchestration will be called the
command will be applied. The command is idempotent so we can run it each
time.
This can be backported for 1.0.3

Closes: https://github.com/rook/rook/issues/2960
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-13 22:15:15 +02:00
Sébastien Han 0317de9096 rgw: change default frontend on nautilus
As per: ceph/ceph#26599, Beast is now the
default fronted for rados gateway.
Newly created cluster as of Nautilus will use it by default.

Re-added version of 03587352d5
Resolves: #2707
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-07 23:10:11 +02:00
Travis Nielsen f2c49ea034 Merge pull request #3165 from d-luu/resource_comparer
ceph: added comparer for resource quantity when checking cluster changes
2019-06-07 11:15:33 -07:00
Sébastien Han 6438d9907e ceph: retry on detecting ceph version
Sometime the kube engine needs a bit of time to return the logs of a
given job and fails to read the stream.
Retrying up to detect the Ceph version seems reasonnable to
overcome this issue.

Fixes: https://github.com/rook/rook/issues/3227
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-07 18:13:38 +02:00
Sébastien Han d3c8d5612b ceph: add resource limit check for rbdmirror
The memory check was missing and will be trigger if resources limit are
configured for the rbdmirror pod.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-07 15:06:16 +02:00
Sébastien Han 24abbddf14 ceph: fix pod memory check
This commit fixes the second test case where limit and request are
either identical or different but still we use limit as a value.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-07 15:06:16 +02:00
DL186038 d22886da37 ceph: added resource quantity comparer when checking cluster changes
the resource struct has non-exportable fields which panics when
doing the diff comparison

Signed-off-by: d-luu <david@davidluu.info>
2019-06-06 15:39:50 -05:00
rohan47 a380ca7dc1 fixed ceph pg dump pgs_brief json deserialization
Signed-off-by: rohan47 <rohgupta@redhat.com>
2019-06-06 21:54:37 +05:30
Travis Nielsen d5523086b6 Merge pull request #3265 from leseb/rgw-liveness
ceph: rgw add liveness probe check
2019-06-06 10:05:27 -06:00
Sébastien Han 241296be8a ceph: rgw add liveness probe check
We now check if the pod responds on http port 80.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-06-05 14:35:15 +02:00
Noah Watkins 71c88ae9e9 ceph: log osd recovery at info level
for the osd recovery case (DOWN->UP) log at a matching log level as the
message indicating the OSD went down.

fixes: #2904

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-06-04 13:57:49 -07:00
Travis Nielsen 1daa70f79c Merge pull request #3085 from christianhuening/feat-2271
added Liveness Probe to Ceph Mgr
2019-05-28 17:26:12 -06:00
Blaine Gardner e8a638dcba ceph nfs: fix the host path for pod volume
Path previously did not use dataDirHostPath. It should, so fix this.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-05-19 06:33:04 -06:00
Christian Hüning f0699c40bf Added Liveness Probe to Ceph Mgr
This will configure a liveness probe for every ceph mgr deployment

Signed-off-by: Christian Hüning <christian.huening@figo.io>
2019-05-17 07:24:28 +02:00
Santosh Pillai 7d5afa29b8 added device class pool property
- Updated code to use deviceClass property when a pool is created for both "replicated" and "erasure code".
- Updated "ceph crush rule create-..." command to use "create-replicated" instead of "create-simple"
- Updated unit tests
- Updated (ceph-pool-crd.md) documentation to reflect the changes.
- Updated pending release notes.

Signed-off-by: Santosh Pillai <sapillai@redhat.com>
2019-05-15 16:45:39 +05:30
Jose A. Rivera f3a1512983 Read node labels for provider topology
Signed-off-by: Jose A. Rivera <jarrpa@redhat.com>
2019-05-14 13:18:00 -05:00
travisn cb82f36eba ceph: no need for --public-bind-addr with host networking
The mons require the --public-bind-addr when host networking is not in use
since the pod ip will be different from the service ip that will be
advertised to clients. When host networking is enabled, there is no
need for this argument since the endpoints will be the same to bind
inside the pod as what is advertised to clients.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-05-10 15:07:46 -06:00
travisn dbe3da97a1 ceph: preserve the non-default port for hostnetworking
Clusters using host networking will not work if the non-default port is
being used by the mons. Previous to rook 1.0 the mons were all
using the non-default port 6790 and now they are using the default port
of 6789 from 1.0. Now host networking will preserve the non-default port
to allow upgrades to continue working.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-05-10 12:24:57 -06:00
Noah Watkins e33eb98047 ceph: run self-signed cert creation with timeout
fixes: #2784

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-05-06 10:10:05 -07:00
Noah Watkins c2dbb93096 ceph: simplify the ceph command execution interface
this adds a structure to hold various settings related to executing ceph
cli commands, and introduces a common interface for configuring and
running such commands.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-05-06 10:10:05 -07:00
Alexander Trost 7cf8e99a8c ceph-op: Remove taint check in "is node scheduable" func
This removes the unnecessary check for taints on nodes in the
`GetNodeSchedulable()` function. This fixes that even though the user
has specified `tolerations` for, e.g., `NoSchedule` taints, that would
be ignored and the node(s) directly be "marked" as unschedulable.

Signed-off-by: Alexander Trost <galexrt@googlemail.com>
2019-05-02 18:05:46 +02:00
Noah Watkins a661b2db02 ceph: rewind http bind fix on upgrade
when upgrading to a newer version of ceph that doesn't need the http
binding fix, this patch unwinds the configuration fix which may still be
present when upgrading from a cluster without the fix. if the
configuration fix remains then the dashboard and prometheus servers will
not be able to properly configure themselves.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-05-01 22:30:12 -07:00
Travis Nielsen 221560bf24 Merge pull request #3090 from kshlm/delete-active-job-if-requested
ceph: Delete and recreate rook-ceph-detect-version job
2019-05-01 10:16:18 -06:00
Travis Nielsen 726efd017a Merge pull request #3089 from travisn/osds-query-latest-nodes
Query the latest list of nodes for osd orchestration
2019-05-01 10:16:04 -06:00
travisn 77c7eb52f9 ceph: query the latest list of nodes for osd orchestration
The list of nodes to be queried for configuring OSDs
should include all of the K8s nodes before the placement
or other filters are applied. The operator was sometimes
missing configuring OSDs on nodes if the discover pod
hadn't finished running on a node. Now the osd prepare
job will be triggered on a node even if the discover pod
hasn't yet run there.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-30 17:07:12 -06:00
travisn 20f6af8a4d remove obsolete flex driver version checks
The flex driver no longer needs to check whether a kubelet
restart is needed or whether the namespace-specific driver
can be supported. The copy of k8s 1.13 types can also
be removed since we have updated to a newer k8s client.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-30 14:50:33 -06:00
Noah Watkins fc7baef6aa ceph: add tests for dashboard http bind fix
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-30 10:57:37 -07:00
Noah Watkins 1efb8af87b ceph: selectively apply http bind addr fix
fixes: #2492

the http bind address fix in rook address an upstream issue in ceph in
which cherrypy break when binding to `::` when the host doesn't have an
ipv6 address. it fixes the issue by always binding prometheus and the
dashboard webservers to the pod address. by doing this, it prevents
using the kubectl port-forward technique, which some users would like.

a patch to cherrypy is upstream in ceph. this PR checks the ceph version
and selectively disables the fix so that both port-forward and the
normal service method works. this is the verison matrix:

          FIX     VER
          ---     ---
luminous merged (>= 12.2.12)
mimic    merged (>= 13.2.6)
nautilus merged (>= 14.1.1)

when all the containers for these versions exists and sufficient time
passes the unreachable code guarded by the version checks in this patch
can be removed.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-30 10:06:39 -07:00
Noah Watkins 3300497c29 ceph: add mimic version check interface
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-30 10:06:39 -07:00
Kaushal M 5e8375e921 ceph: Delete and recreate rook-ceph-detect-version job
The ceph controller will delete and recreate the detec-version job, when
an already existing job is found. This helps avoid cluster creation from
getting hung up on waiting for an already existing/stale job to finish.

Fixes #3081.

Signed-off-by: Kaushal M <kshlmster@gmail.com>
2019-04-30 18:19:56 +05:30
wangxiao86 a5a8d22478 Set the module's server_addr path with the mgr's name.
To enable multiple mgrs to run properly, we need to set the respective server_addr path with the mgr's name for the modules of each mgr.

When multiple mgrs are running, the modules( e.g. dashboard and Prometheus ) of each mgr need to bind its own server_addr when it is started. The way to get server_addr when ceph mgr starts the module is as follows:  first,  get the server_addr value by "mgr/<module-name>/<mgr-daemonid>/server_addr" path;  and if this path does not exist, then get the server_addr value by" mgr/<module-name>/server_addr" path.
According to the original rook setting for server_addr,  the modules of all mgrs take the same server_addr on ceph luminous and mimic, this produces a wrong result that only one mgr's modules can run properly ( the modules of other mgrs failed to start with the wrong server_addr).
This change is made to be compatible with ceph luminous, mimic, and nautilus.  and Modify according to: the "HOST NAME AND PORT" description of http://docs.ceph.com/docs/master/mgr/dashboard/ .

Signed-off-by: wangxiao86 <wangxiao1@sensetime.com>
2019-04-30 17:14:07 +08:00
Noah Watkins 72b3f79ce7 ceph: skip orchestration when cluster is not yet configured
this skips orchestration when the device configmap receives an update
but the cluster hasn't been fully configured, which may leave some state
undefined and cause nil pointer dereferences accessing the cluster info
structure.

fixes: #3001

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-26 11:23:35 -07:00
Travis Nielsen 5f466b903e Merge pull request #3065 from noahdesu/hotplug-false-postives
ceph: check for equiv device config maps in operator
2019-04-26 12:03:04 -06:00
Noah Watkins fbc64ff8f4 ceph: check for equiv device config maps in operator
this is an attempt to reduce the number of false positive config map
updates used to trigger orchestration for hotplugging.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-26 09:37:25 -07:00
travisn b6568b07b4 ceph: stop disabling scrubbing during osd orchestration
Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-26 09:35:50 -06:00
Sébastien Han 5f16f8e27f ceph: fix upgrade for monitor listening on 6790 port
This commits fixes the upgrade for cluster running on 0.9.x and thus
have monitors listening on 6790.

The only downside is that the new messengers 2 protocol can not be
enabled because Ceph expects monitors to listen on 6789.
In order to enable messengers 2 your monitors will have to failover so
they can pick up both 6789 and 3300 ports, then and only then messenger
2 can be enabled.

Fixes: https://github.com/rook/rook/issues/3048 and
https://github.com/rook/rook/issues/3049
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-04-25 16:44:38 +02:00
Travis Nielsen b70bec1a76 Merge pull request #3036 from rohantmp/revertMainMon
Revert "Merge pull request #2986 from rohantmp/mainMon"
2019-04-24 15:02:58 -06:00
Travis Nielsen d11466fced Merge pull request #2900 from mackIOConsulting/fix-mgr-settings
Fix mgr settings
2019-04-24 14:44:20 -06:00
Rohan CJ fb4d23345e Revert "Merge pull request #2986 from rohantmp/mainMon"
This reverts commit b7a564e929, reversing
changes made to 6996571b5c.

Reverting: Linking ceph noout to a node taint.
The taint is applied for reasons other than maintenance as well.
This also needs a way to respect user-set noout.
Need to revisit this with a better design.

Signed-off-by: Rohan CJ <rohantmp@gmail.com>
2019-04-24 22:38:38 +05:30
Sébastien Han ab439f6602 Ceph: fix upgrade to Nautilus
This commits allows the upgrade from Mimic to Nautilus to work by:

* removing the v2 brackets on the operator ceph config file generation,
we stick with v1 and can revert this back once we deprecate mimic there
is no rush since mons keep on listening to v1.

* enable messengers 2 when the cluster runs on Nautilus

Fixes: https://github.com/rook/rook/issues/2973
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-04-24 10:57:29 +02:00
Maximilian Mack bdbd67fd7f ceph: fix mgr - dashboard settings
This fix passes on the mgrName to the configureDashboard function.

Signed-off-by: Maximilian Mack <max@mack.io>
2019-04-24 08:26:22 +02:00
Travis Nielsen fe2a291918 Merge pull request #3032 from noahdesu/hotplug-ignore-rbd
Hotplug ignore rbd
2019-04-23 17:47:08 -06:00
Noah Watkins 515b02deec ceph: add debug logging to operator hotplug action
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-23 13:44:30 -07:00
Noah Watkins f2fc644190 ceph: add operator option to disable hotplug
this adds a new operator environment variable that allows automatic
orchestration triggered by device discovery to be disabled.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-04-23 13:44:30 -07:00
travisn 40697eacbf ceph: update the nfs ganesha deployment if it already exists
Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-23 13:36:45 -06:00
travisn 9dcddb5bed ceph: upgrade all daemons during ceph upgrade
The CephCluster CR contains settings that are needed by other
CRs to configure the Ceph daemons. When the CephCluster CR
is updated, the updates will now be passed on to each of the
CR controllers to ensure the daemons are updated properly
without requiring an operator restart.

When calling the controllers from another controller,
we ensure that only a single goroutine is handling
CRs at any given time to prevent contention across
multiple CRs of the same type.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-23 13:36:45 -06:00
Louis f88beafc56 update ROOK_MON_OUT_TIMEOUT to 600s
Signed-off-by: Louis <gofaceme@gmail.com>
2019-04-23 11:19:23 +08:00
Blaine Gardner f4e2be50a6 ceph spec: add ceph-version label to controllers
Add a 'ceph-version' label to application controllers in the same manner
as the prior 'rook-version' label to help with upgrades.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-04-18 15:09:33 -06:00
Travis Nielsen 9f126e3902 Merge pull request #3009 from SUSE/mon-remove-legacy-replicaset-deletion
ceph mon: remove code to delete legacy replicasets
2019-04-18 10:35:23 -06:00
travisn bc87b440fb update to the k8s 1.14 client libraries
Signed-off-by: travisn <tnielsen@redhat.com>
2019-04-18 07:51:41 -06:00