When OSDs are running on Nautilus we always disable old osd features and
aplpy the onces for Nautilus as described in the upgrade doc.
During an upgrade or the next time an orchestration will be called the
command will be applied. The command is idempotent so we can run it each
time.
This can be backported for 1.0.3
Closes: https://github.com/rook/rook/issues/2960
Signed-off-by: Sébastien Han <seb@redhat.com>
As per: ceph/ceph#26599, Beast is now the
default fronted for rados gateway.
Newly created cluster as of Nautilus will use it by default.
Re-added version of 03587352d5Resolves: #2707
Signed-off-by: Sébastien Han <seb@redhat.com>
Sometime the kube engine needs a bit of time to return the logs of a
given job and fails to read the stream.
Retrying up to detect the Ceph version seems reasonnable to
overcome this issue.
Fixes: https://github.com/rook/rook/issues/3227
Signed-off-by: Sébastien Han <seb@redhat.com>
The memory check was missing and will be trigger if resources limit are
configured for the rbdmirror pod.
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit fixes the second test case where limit and request are
either identical or different but still we use limit as a value.
Signed-off-by: Sébastien Han <seb@redhat.com>
for the osd recovery case (DOWN->UP) log at a matching log level as the
message indicating the OSD went down.
fixes: #2904
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
- Updated code to use deviceClass property when a pool is created for both "replicated" and "erasure code".
- Updated "ceph crush rule create-..." command to use "create-replicated" instead of "create-simple"
- Updated unit tests
- Updated (ceph-pool-crd.md) documentation to reflect the changes.
- Updated pending release notes.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
The mons require the --public-bind-addr when host networking is not in use
since the pod ip will be different from the service ip that will be
advertised to clients. When host networking is enabled, there is no
need for this argument since the endpoints will be the same to bind
inside the pod as what is advertised to clients.
Signed-off-by: travisn <tnielsen@redhat.com>
Clusters using host networking will not work if the non-default port is
being used by the mons. Previous to rook 1.0 the mons were all
using the non-default port 6790 and now they are using the default port
of 6789 from 1.0. Now host networking will preserve the non-default port
to allow upgrades to continue working.
Signed-off-by: travisn <tnielsen@redhat.com>
this adds a structure to hold various settings related to executing ceph
cli commands, and introduces a common interface for configuring and
running such commands.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
This removes the unnecessary check for taints on nodes in the
`GetNodeSchedulable()` function. This fixes that even though the user
has specified `tolerations` for, e.g., `NoSchedule` taints, that would
be ignored and the node(s) directly be "marked" as unschedulable.
Signed-off-by: Alexander Trost <galexrt@googlemail.com>
when upgrading to a newer version of ceph that doesn't need the http
binding fix, this patch unwinds the configuration fix which may still be
present when upgrading from a cluster without the fix. if the
configuration fix remains then the dashboard and prometheus servers will
not be able to properly configure themselves.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
The list of nodes to be queried for configuring OSDs
should include all of the K8s nodes before the placement
or other filters are applied. The operator was sometimes
missing configuring OSDs on nodes if the discover pod
hadn't finished running on a node. Now the osd prepare
job will be triggered on a node even if the discover pod
hasn't yet run there.
Signed-off-by: travisn <tnielsen@redhat.com>
The flex driver no longer needs to check whether a kubelet
restart is needed or whether the namespace-specific driver
can be supported. The copy of k8s 1.13 types can also
be removed since we have updated to a newer k8s client.
Signed-off-by: travisn <tnielsen@redhat.com>
fixes: #2492
the http bind address fix in rook address an upstream issue in ceph in
which cherrypy break when binding to `::` when the host doesn't have an
ipv6 address. it fixes the issue by always binding prometheus and the
dashboard webservers to the pod address. by doing this, it prevents
using the kubectl port-forward technique, which some users would like.
a patch to cherrypy is upstream in ceph. this PR checks the ceph version
and selectively disables the fix so that both port-forward and the
normal service method works. this is the verison matrix:
FIX VER
--- ---
luminous merged (>= 12.2.12)
mimic merged (>= 13.2.6)
nautilus merged (>= 14.1.1)
when all the containers for these versions exists and sufficient time
passes the unreachable code guarded by the version checks in this patch
can be removed.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
The ceph controller will delete and recreate the detec-version job, when
an already existing job is found. This helps avoid cluster creation from
getting hung up on waiting for an already existing/stale job to finish.
Fixes#3081.
Signed-off-by: Kaushal M <kshlmster@gmail.com>
To enable multiple mgrs to run properly, we need to set the respective server_addr path with the mgr's name for the modules of each mgr.
When multiple mgrs are running, the modules( e.g. dashboard and Prometheus ) of each mgr need to bind its own server_addr when it is started. The way to get server_addr when ceph mgr starts the module is as follows: first, get the server_addr value by "mgr/<module-name>/<mgr-daemonid>/server_addr" path; and if this path does not exist, then get the server_addr value by" mgr/<module-name>/server_addr" path.
According to the original rook setting for server_addr, the modules of all mgrs take the same server_addr on ceph luminous and mimic, this produces a wrong result that only one mgr's modules can run properly ( the modules of other mgrs failed to start with the wrong server_addr).
This change is made to be compatible with ceph luminous, mimic, and nautilus. and Modify according to: the "HOST NAME AND PORT" description of http://docs.ceph.com/docs/master/mgr/dashboard/ .
Signed-off-by: wangxiao86 <wangxiao1@sensetime.com>
this skips orchestration when the device configmap receives an update
but the cluster hasn't been fully configured, which may leave some state
undefined and cause nil pointer dereferences accessing the cluster info
structure.
fixes: #3001
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
this is an attempt to reduce the number of false positive config map
updates used to trigger orchestration for hotplugging.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
Signed-off-by: Ashish Ranjan <ashishranjan738@gmail.com>
This commit removes one to one mapping of NFS provisioner and adds a dedicated provisioner for nfs.