Log file path on host for OSDs was dataDirHostPath/log/<namespace>
instead of dataDirHostPath/<namespace>/log as with other daemons.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
PopulateDeviceInfo() in pkg/clusterd/disk.go returns nil as *sys.LocalDisk
if an error has occurred. This causes a nil pointer exception at
getAvailableDevices() in pkg/daemon/ceph/osd/daemon.go. At least the returned
value should be checked at the caller.
I added an explicit error to the return value of PopulateDeviceInfo(). This
naturally revokes an error check at the caller.
I modified PopulateDeviceUdevInfo() in the same file too.
Signed-off-by: morimoto-cybozu <kenji_morimoto@cybozu.co.jp>
The cluster finalizer in some scenarios was not being removed
during cluster removal. Specifically, if the cluster CR had been
modified, the operator would always fail to remove the finalizer.
This could occur when the CR status is updated around the same time
that the cluster is deleted. Therefore, we need to remove the finalizer
from a freshly retrieved instance of the cluster.
Signed-off-by: travisn <tnielsen@redhat.com>
The owner references should be set based on the name, otherwise the
garbage collector will delete the resources at some point.
These references to yugabyte, minio, and cockroach were
missed in the original fix for these references.
Signed-off-by: travisn <tnielsen@redhat.com>
The resources for rgw, mds, and nfs were being cleaned up explicitly when
the CR was deleted. This is legacy from before the owner references were
set on the CRs. Now the owner references are set to the appropriate CRs so
the resources will be cleaned up upon deletion of the specific CR
rather than the cluster CR.
Signed-off-by: travisn <tnielsen@redhat.com>
The full path to vgchange should not be necessary. For some
downstream releases the path may even be in a different path.
Therefore we require that vgchange be in the path.
Signed-off-by: travisn <tnielsen@redhat.com>
The access of user will be revoked by Droppolicy(). And user unlink
need to perform only if it is the owner. For Delete() it makes sense
to call user unlink, but it is not required for Revoke()
Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
Following corrections are needed for bucket policy to work properly
* Change Principal:{"AWS":["<username>"]} to Principal:{"AWS":["arn:aws:iam:::user/<username>"]}
* Add missing "PutObject" policy in Action[]
* Add <bucket/*> to Resource[] for accessing contents inside the bucket
Signed-off-by: Jiffin Tony Thottan <jthottan@redhat.com>
We now differentiate the cases where:
* we only consume the external cluster
* we consume the external cluster as well as creating stateless
resources in Kubernetes (bootstrap mds,rgw, nfs)
This is mostly controlled via the image property spec. If not defined,
not extra CRs won't be able to be created.
Now the external cluster feature supports Ceph cluster as of Luminous 12.2.
Signed-off-by: Sébastien Han <seb@redhat.com>
Update pkg/operator/ceph/cluster/osd/spec.go
Rook v1.1.2 uses a different command to start an OSD, and seems to
have problems when --osd-memory-target contains a number with a
decimal point.
Co-Authored-By: Sébastien Han <seb@redhat.com>
Signed-off-by: Erik Knudtson <eknudtson@discogsinc.com>
When osd nodes do not finish the prepare jobs, the user remove the
nodes manully and it causes nodes always is
'OrchestrationStatusStarting' status.
Try to remove them after completeProvisionTimeout.
Signed-off-by: James Lu <jamesluhz@gmail.com>
Set the default Ceph CSI images as vars in the code instead of consts.
This allows these values to be overridden at build time with Go linker
-X flags. This allows users to build Rook in such a way that it will
automatically update the CSI images to a custom opinionated default
without having to manually manage the environment variable-based
overrides at upgrade time.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
A new CRD property `PreservePoolsOnDelete` has been added to Filesystem(fs) and
Object Store(os) resources in order to increase protection against data loss.
If it is set to `true`, associated pools won't be deleted when the main
resource(fs/os) is deleted. Creating again the deleted fs/os with the same name
will reuse the preserved pools.
Signed-off-by: Juan Miguel Olmo Martínez <jolmomar@redhat.com>
provided ENV variables to configure the grpc
and liveness metrics port for both cephfs and
rbd CSI drivers.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Add support for setting Toleration and NodeAffinity
to CSI Provisioner deployment or statefulset and
plugin daemonset through ENV variables.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
function AddNodeAffinity was not adding any node
affinity instead it was forming the nodeaffinity
object. renamed it to GenerateNodeAffinity for more
meaningful
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Add a finalizer to the CephCluster to check
the PV created by CSI so that we can block the
cluster deletion till the PVC are deleted.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
We don't need to force any Ceph config file for ceph-volume based OSDs
since the 'ceph-volume activate' command mount everything at start time
in the right place.
So now, OSD will keep running normally and they will also read back
normally the default location of the ceph configuration:
/etc/ceph/ceph.conf which will allow OSDs to read any configuration
overrides.
We cannot remove the generation of that file in /var/lib/rook/osd* since
we cannot tell at the stage that triggers the generation if we run
ceph-volume or not. So let's keep it around for some time.
If you are wondering how the OSD connects to the mons this is done via
the CEPH_ARGS env variable which was introduced in
https://github.com/rook/rook/pull/3894.
Closes: https://github.com/rook/rook/issues/3926
Signed-off-by: Sébastien Han <seb@redhat.com>
currently we are not checking the pool is empty or
not before deleting, This PR adds a check to check
if any images/snapshots using the pool, if yes it will
not delete the pool
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
A job is started to detect the Ceph version. This job now allows setting
the node affinity and tolerations with the same setting that is
specified for the mon placement. No new placement spec is required, but
we will simply use the same spec that the mons use.
Signed-off-by: travisn <tnielsen@redhat.com>
Previously, we were applying lvm config changes to OSD running on PV. So
now, we do this for any type of cluster deployment (bare metal or
Cloud).
Signed-off-by: Sébastien Han <seb@redhat.com>
Previously the debug flag was turned on for the 'activate' call only,
now 'prepare' will also take advantage of it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Previously, "run dir" was used to place a number of file and config on
dataDirHostPath. Now, since a lot of configs and options have moved
either on the OSD's startup CLI line or remove with bluestore we don't
need to change the default.
Also, it was confusing for user and difficult to find the socket.
The is is not breaking any config in dataDirHostPath since all the
elements (keyrings) are hardcoded in the config file.
Closes: https://github.com/rook/rook/issues/3966
Signed-off-by: Sébastien Han <seb@redhat.com>
When the flex driver is disabled, the check for
volume attachments will fail and the finalizer
is never removed. To avoid this, just log the
failure to list volumes and remove the
finalizer anyway.
Resolves#3912
Signed-off-by: Kristoffer Grönlund <kgronlund@suse.com>
previously we are not updating the service
if it's already present. with this change,
we will update the service. this does not
trigger any error message if service
already present.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
Encrypted OSDs are failing because ceph-volume is unable to determine
how to contact the monitors. Setting the CEPH_CONF env variable for OSD
pods fixes this issue. (#3846)
Signed-off-by: Michael Vollman <michael.b.vollman@gmail.com>
RHEL8 introduces some lvm changes. For OSDs on PVs to work
we need to set allow_changes_with_duplicate_pvs = 1.
This is required since we're copying the block image
under another location.
Signed-off-by: travisn <tnielsen@redhat.com>