Commit Graph
1353 Commits
Author SHA1 Message Date
Blaine Gardner 3584e8ccce ceph: osd: fix wrong log file path on host
Log file path on host for OSDs was dataDirHostPath/log/<namespace>
instead of dataDirHostPath/<namespace>/log as with other daemons.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-10-15 10:10:06 -06:00
Travis Nielsen fdd92ec883 Merge pull request #4100 from morimoto-cybozu/clarify-errors-in-clusterd
Return explicit errors
2019-10-15 08:57:06 -06:00
morimoto-cybozu 815927d943 clusterd: return explicit errors
PopulateDeviceInfo() in pkg/clusterd/disk.go returns nil as *sys.LocalDisk
if an error has occurred.  This causes a nil pointer exception at
getAvailableDevices() in pkg/daemon/ceph/osd/daemon.go.  At least the returned
value should be checked at the caller.
I added an explicit error to the return value of PopulateDeviceInfo().  This
naturally revokes an error check at the caller.
I modified PopulateDeviceUdevInfo() in the same file too.

Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
2019-10-15 08:30:45 +00:00
travisn bc68440006 ceph: more robust removal of cluster finalizer
The cluster finalizer in some scenarios was not being removed
during cluster removal. Specifically, if the cluster CR had been
modified, the operator would always fail to remove the finalizer.
This could occur when the CR status is updated around the same time
that the cluster is deleted. Therefore, we need to remove the finalizer
from a freshly retrieved instance of the cluster.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-14 09:29:46 -06:00
travisn b0a3bea421 ownerrefs: set owners based on name instead of namespace
The owner references should be set based on the name, otherwise the
garbage collector will delete the resources at some point.
These references to yugabyte, minio, and cockroach were
missed in the original fix for these references.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-14 09:29:45 -06:00
travisn 5db209fc5b ceph: allow resources to be deleted via owner references
The resources for rgw, mds, and nfs were being cleaned up explicitly when
the CR was deleted. This is legacy from before the owner references were
set on the CRs. Now the owner references are set to the appropriate CRs so
the resources will be cleaned up upon deletion of the specific CR
rather than the cluster CR.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-14 09:29:45 -06:00
travisn 22391377a8 ceph: remove full path from vgchange call
The full path to vgchange should not be necessary. For some
downstream releases the path may even be in a different path.
Therefore we require that vgchange be in the path.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-10-14 09:29:45 -06:00
Sébastien Han d7956facdb Merge pull request #4075 from rohantmp/mdsFix
[Ceph] Fix: PDB for MDS not created when activeStandby:true and activeCount: 1
2019-10-11 11:02:12 +02:00
Rohan CJ 0d9442982b ceph: Fix PDB for MDS is not created if minAvailable <= 1.
It should be minAvailable < 1.

Signed-off-by: Rohan CJ <rohantmp@gmail.com>
2019-10-10 21:34:20 +05:30
Rohan CJ 3c5856fc2a Ceph: Take activeStandby into account for CephFileSystem disruption budget.
When activeStandby is set on the CephFilesystem, the effective
activeCount will be incremented.

Signed-off-by: Rohan CJ <rohantmp@gmail.com>
2019-10-10 21:31:58 +05:30
Jiffin Tony Thottan 3f3af19a13 objectstore: remove user unlink in Revoke()
The access of user will be revoked by Droppolicy(). And user unlink
need to perform only if it is the owner. For Delete() it makes sense
to call user unlink, but it is not required for Revoke()

Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
2019-10-10 19:21:45 +05:30
Jiffin Tony Thottan 69e4144b0f objectstore: modify existing bucket policy for brownfield cases
Following corrections are needed for bucket policy to work properly

* Change Principal:{"AWS":["<username>"]} to Principal:{"AWS":["arn:aws:iam:::user/<username>"]}
* Add missing "PutObject" policy in Action[]
* Add <bucket/*> to Resource[] for accessing contents inside the bucket

Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
2019-10-10 19:21:17 +05:30
Sébastien Han 28a898d9f3 Merge pull request #4046 from eknudtson/osd-memory-target-fix
remove decimal point for osdMemoryTargetValue
2019-10-10 09:27:50 +02:00
Sébastien Han 148f8da1fb ceph: relax pre-requisite for external cluster
We now differentiate the cases where:

* we only consume the external cluster
* we consume the external cluster as well as creating stateless
resources in Kubernetes (bootstrap mds,rgw, nfs)

This is mostly controlled via the image property spec. If not defined,
not extra CRs won't be able to be created.

Now the external cluster feature supports Ceph cluster as of Luminous 12.2.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-10-08 10:52:01 +02:00
Erik KnudtsonandSébastien Han ec4d0c8925 ceph: remove decimal places for osdMemoryTargetValue
Update pkg/operator/ceph/cluster/osd/spec.go

Rook v1.1.2 uses a different command to start an OSD, and seems to
have problems when --osd-memory-target contains a number with a
decimal point.

Co-Authored-By: Sébastien Han <seb@redhat.com>
Signed-off-by: Erik Knudtson <eknudtson@discogsinc.com>
2019-10-07 09:32:49 -07:00
Sébastien Han 233df963e7 ceph: fix various spelling issues
Correct spelling.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-10-03 11:42:40 +02:00
James Lu b8d7382b7a Ceph: nodes always in OrchestrationStatusStarting
When osd nodes do not finish the prepare jobs, the user remove the
nodes manully and it causes nodes always is
'OrchestrationStatusStarting' status.
Try to remove them after completeProvisionTimeout.

Signed-off-by: James Lu <jamesluhz@gmail.com>
2019-10-03 08:08:46 +08:00
Sébastien Han 418cbb9464 Merge pull request #4018 from SUSE/var-csi-images
ceph: set default Ceph CSI images as var not const
2019-10-01 09:39:06 +02:00
Blaine Gardner 24d2187f1a ceph: set default Ceph CSI images as var not const
Set the default Ceph CSI images as vars in the code instead of consts.
This allows these values to be overridden at build time with Go linker
-X flags. This allows users to build Rook in such a way that it will
automatically update the CSI images to a custom opinionated default
without having to manually manage the environment variable-based
overrides at upgrade time.

Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
2019-09-30 12:15:32 -06:00
Juan Miguel Olmo Martínez 2de0787fc8 New setting to avoid delete pools on <fs>/<os> resources deletion
A new CRD property `PreservePoolsOnDelete` has been added to Filesystem(fs) and
Object Store(os) resources in order to increase protection against data loss.
If it is set to `true`, associated pools won't be deleted when the main
resource(fs/os) is deleted. Creating again the deleted fs/os with the same name
 will reuse the preserved pools.

Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
2019-09-30 08:13:39 +02:00
Travis Nielsen 91b8a96fb6 Merge pull request #3994 from leseb/rm-osd-config
Ceph: fix rook-config-overrides for OSDs
2019-09-27 10:07:41 -06:00
Madhu Rajanna f54099df0c CSI: Make metrics and liveness port configurable
provided ENV variables to configure the grpc
and liveness metrics port for both cephfs and
rbd CSI drivers.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-09-27 17:31:45 +05:30
Madhu Rajanna 9d298fe4f0 Add Toleration and NodeAffinity to CSI
Add support for setting Toleration and NodeAffinity
to CSI Provisioner deployment or statefulset and
plugin daemonset through ENV variables.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-09-27 10:23:12 +05:30
Madhu Rajanna 55e9340737 rename kserrors to k8serrors
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-09-27 10:05:45 +05:30
Madhu Rajanna 3af3cd63ab Rename AddNodeAffinity to GenerateNodeAffinity
function AddNodeAffinity was not adding any node
affinity instead it was forming the nodeaffinity
object. renamed it to GenerateNodeAffinity for more
meaningful

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-09-27 10:05:45 +05:30
Dmitry Yusupov 37a9e6b0f0 Merge pull request #3995 from SergeyKaydalov/edgefs-kvssd
EdgeFS: support for rtkvs disk engine and Samsung KVSSD
2019-09-26 10:34:16 -07:00
Travis Nielsen 4d221288ce Merge pull request #3967 from Madhu-1/fix-3913
Add check for csi volumes before cluster delete
2019-09-26 11:13:07 -06:00
Travis Nielsen 11348e8b67 Merge pull request #3972 from Madhu-1/pool-delete
Check Pools is in use before deleting it
2019-09-26 11:12:05 -06:00
Sergey Kaydalov 07ba73bb8f EdgeFS: support for rtkvs disk engine and Samsung KVSSD
Signed-off-by: Sergey Kaydalov <kayserg@gmail.com>
2019-09-26 18:02:32 +03:00
Madhu Rajanna ec9784203e Add check for csi volumes before cluster delete
Add a finalizer to the CephCluster to check
the PV created by CSI so that we can block the
cluster deletion till the PVC are deleted.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-09-26 20:15:11 +05:30
Sébastien Han ae8e45585f ceph: osd add fsid to daemon flag
Now that the ceph.conf is gone, it's safer to run the OSD with the
cluster fsid flag.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-09-26 15:29:24 +02:00
Sébastien Han a4a180b8fb ceph: osd remove config file
We don't need to force any Ceph config file for ceph-volume based OSDs
since the 'ceph-volume activate' command mount everything at start time
in the right place.
So now, OSD will keep running normally and they will also read back
normally the default location of the ceph configuration:
/etc/ceph/ceph.conf which will allow OSDs to read any configuration
overrides.
We cannot remove the generation of that file in /var/lib/rook/osd* since
we cannot tell at the stage that triggers the generation if we run
ceph-volume or not. So let's keep it around for some time.
If you are wondering how the OSD connects to the mons this is done via
the CEPH_ARGS env variable which was introduced in
https://github.com/rook/rook/pull/3894.

Closes: https://github.com/rook/rook/issues/3926
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-09-26 15:07:53 +02:00
Madhu Rajanna 4efba0247d Check Pools is in use before deleting it
currently we are not checking the pool is empty or
not before deleting, This PR adds a check to check
if any images/snapshots using the pool, if yes it will
not delete the pool

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-09-26 11:41:47 +05:30
travisn 0f5f2a81d7 ceph: allow setting affinity on the ceph version job
A job is started to detect the Ceph version. This job now allows setting
the node affinity and tolerations with the same setting that is
specified for the mon placement. No new placement spec is required, but
we will simply use the same spec that the mons use.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-09-25 10:10:45 -06:00
Travis Nielsen b99f6efb03 Merge pull request #3979 from Madhu-1/change-csi-tag
use new cephcsi release
2019-09-25 06:52:45 -06:00
Madhu Rajanna 0a516d5f32 use v1.2.1 cephcsi release
update CSI deployment templates and
doc to use new cephcsi v1.2.1 release

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-09-25 16:08:55 +05:30
Sébastien Han a42cc44d11 Merge pull request #3894 from mvollman/osd_ceph_conf
ceph: fix encrypted osd startup
2019-09-25 09:00:06 +02:00
Sébastien Han 24fc94cc57 Merge pull request #3955 from leseb/cv-debug-prepare
ceph: various osd fixes and improvements with lvm and ceph-volume
2019-09-25 08:46:34 +02:00
Sébastien Han 21e44c29e9 Merge pull request #3969 from leseb/osd-socket
ceph: osd reset 'run dir' to default location
2019-09-25 08:45:46 +02:00
Sébastien Han 42cebf6311 ceph: osd run lvm config all the time
Previously, we were applying lvm config changes to OSD running on PV. So
now, we do this for any type of cluster deployment (bare metal or
Cloud).

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-09-24 22:49:39 +02:00
Sébastien Han 3404460d3c ceph: osd, fix lvm filter regexp
The lvm documentation uses commas to separate the regular expressions.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-09-24 17:17:29 +02:00
Sébastien Han 9b2afbca14 ceph: osd add debug logs for ceph-volume on prepare
Previously the debug flag was turned on for the 'activate' call only,
now 'prepare' will also take advantage of it.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-09-24 17:17:29 +02:00
Travis Nielsen 9affd75752 Merge pull request #3959 from Madhu-1/fix-svc-crt
Update service if already present
2019-09-24 08:44:14 -06:00
Sébastien Han ad514606bf ceph: osd reset 'run dir' to default location
Previously, "run dir" was used to place a number of file and config on
dataDirHostPath. Now, since a lot of configs and options have moved
either on the OSD's startup CLI line or remove with bluestore we don't
need to change the default.

Also, it was confusing for user and difficult to find the socket.
The is is not breaking any config in dataDirHostPath since all the
elements (keyrings) are hardcoded in the config file.

Closes: https://github.com/rook/rook/issues/3966
Signed-off-by: Sébastien Han <seb@redhat.com>
2019-09-24 10:27:59 +02:00
Sébastien Han d8813513eb ceph: fix various spelling issues
Correct spelling.

Signed-off-by: Sébastien Han <seb@redhat.com>
2019-09-24 10:26:22 +02:00
Kristoffer Grönlund fb599dd02b Ceph: Remove finalizer even if flex is disabled
When the flex driver is disabled, the check for
volume attachments will fail and the finalizer
is never removed. To avoid this, just log the
failure to list volumes and remove the
finalizer anyway.

Resolves #3912

Signed-off-by: Kristoffer Grönlund <kgronlund@suse.com>
2019-09-24 08:22:30 +02:00
Madhu Rajanna bdfdac8ee4 Update service if already present
previously we are not updating the service
if it's already present. with this change,
we will update the service. this does not
trigger any error message if service
already present.

Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
2019-09-24 09:43:02 +05:30
Michael Vollman e12bbd6f7c ceph: fix encrypted osd startup
Encrypted OSDs are failing because ceph-volume is unable to determine
how to contact the monitors.  Setting the CEPH_CONF env variable for OSD
pods fixes this issue. (#3846)

Signed-off-by: Michael Vollman <michael.b.vollman@gmail.com>
2019-09-20 16:36:52 -04:00
travisn f3ce214822 ceph: configure additional lvm settings for rhel8
RHEL8 introduces some lvm changes. For OSDs on PVs to work
we need to set allow_changes_with_duplicate_pvs = 1.
This is required since we're copying the block image
under another location.

Signed-off-by: travisn <tnielsen@redhat.com>
2019-09-20 11:51:07 -06:00
Travis Nielsen 5f37de8abc Merge pull request #3927 from Madhu-1/fix-3921
Make kubelet path configurable in operator for csi
2019-09-20 10:45:27 -06:00