defines a new entry point "close-encrypted-devices" in osd job.
The job is used by osd-replacement flow to close dm-crypt mappings
on host for destroyed encrypted OSDs. The job is owned by OSD
health goroutine implementing destroy phase of osd-replacement.
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
This change refactors the getSecret method somewhat toward cleaner code.
The only additional change in behavior is the introduction of a switch instead
of multiple if's and error treatment of the case of an empty secret.
Signed-off-by: subhamkrai <srai@redhat.com>
Allow force creation of OSD on a device previously configured in a different ceph cluster.
The OSD prepare pod job will zap the OSD disk using ceph-volume lvm zap
if that disk has metadata from a different ceph cluster.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
Implements #14733. Allows to set IDs of external mons to
Cluster CRD. Rook will not remove external mons from quorum
and will add external mon addresses to mon endpoints.
Use-case for external mon is to maintain quorum for 2-AZ
k8s cluster in case of zone outage.
Signed-off-by: Artem Torubarov <artem.torubarov@clyso.com>
This is similar to #14052 we did for radosnamespace
and this is an extension to support cleanup
at the blockpool level to cleanup the images
and the snapshots in a pool.
Signed-off-by: Madhu Rajanna <madhupr007@gmail.com>
The default service account isn't passed to the multus validation test
when no config file is used. Fix this.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
In order to help users check that they have implemented the newly-added
Multus host configuration prerequisites, add a check to the validation
tool to verify connectivity.
Because users who are already running clusters with Multus enabled, add
a flag that allows users to only check for host configuration
prerequisites. This mode will not start the large number of clients that
would normally be started because those clients could disrupt a running
Rook cluster negatively.
Host checking pods require host network access. Many Kubernetes
distributions have pod security features enabled. In order to allow
non-Vanilla distros to run this tool, allow specifying a service account
that pods will run as, which can be configured by the admin to allow
test pods.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
fix UpdateActiveMgrLabel to retry label update on failure. Previously,
the function always returned the current manager, even if update
failed, leading to inconsistent labels. This fix ensures retries on
failures.
Fixes: https://github.com/rook/rook/issues/13601
Signed-off-by: Redouane Kachach <rkachach@ibm.com>
Cleanup the resources created by subvolumegroup
when its deleted. Following resources will be cleaned up:
- OMAP value
- OMAP keys
- Clones
- Snapshots
- Subvolumes
Signed-off-by: sp98 <sapillai@redhat.com>
Go is not cgroup aware and by default will set GOMAXPROCS to the number
of available threads, regardless of whether it is within the allocated
quota. This behaviour causes high amount of CPU throttling and degraded
application performance.
Fixes: #13815
Signed-off-by: Thomas Way <thomas@6f.io>
create the ceph config and keyring file in the osd prepare pod
before starting the OSD migration. These files are needed to
run any ceph command.
Signed-off-by: sp98 <sapillai@redhat.com>
Adding callback function in the osd removal method
as in downstream there is requirement of adding extra
check before proceeding with osd removal.
Signed-off-by: subhamkrai <srai@redhat.com>
Add the ability to specify node profiles in the multus validation test.
This addresses a few points of early feedback on the validation tool.
Statements below critique the tool's behavior before this patch.
1. The tool assumes all daemons are on public and cluster network, which
means users who have a significantly smaller cluster net (a
design choice) cannot run a single test to determine if Rook is
likely to install correctly.
2. The tool does not have placement options to select only a subset of
Kubernetes nodes to run validation on.
3. Users of multus seem to have a dedicated pool of storage nodes more
often than the average Rook install. This makes sense for security-
and perforance-minded users. The tool cannot run a single test to
verify storage-only and general-workload nodes at one time.
These points are addressed by allowing users to specify configurations
for different "NodeTypes."
Each NodeType config has options for selecting the number of OSDs as
well as the number of other (non-OSD) Ceph daemons. This limits the
unnecessary exhaustion of cluster network addresses from critique 1.
Each NodeType config has its own placement (critique 2).
Users can define as many NodeTypes as needed to test the network for
their planned CephCluster. Specifically, this allows the tool to test
storage-only nodes and generalized-workload nodes at the same time. An
arbitrary number of NodeTypes are allowed to support even more highly
specialized cluster setups, such as multiple tiers of storage nodes
where some storage-only nodes may run more OSDs than others.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
If osd store is updated in the ceph cluster, then
delete OSDs one by one, cleanup disks and provision a new OSD on
the same disk
Signed-off-by: sp98 <sapillai@redhat.com>
The function ParseMonEndpoints existed twice, once in the mon package
and once in the controller package. It has now been deleted from the
mon package. Usages in the mon package now reference the function
in the controller package.
Signed-off-by: Henry Davies <henrydavies@hotmail.co.uk>
The previous client loader routine assumed KUBECONFIG would be set in
CLI environments. Instead, now take this approach:
1. If KUBECONFIG is set, that is the de-facto override that informs the
tool it is being run in a CLI environment.
2. Otherwise, try creating a client from the default kube config file.
3. If that fails, assume the tool is running in a Kubernetes Pod.
This will continue supporting dev/test environments so that building the
rook container is not necessary for development. It also allows support
for highly opinionated environments (like deployed by OpenShift CSVs)
where it's not possible to install the Rook operator without also
deploying a CephCluster.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
in the existing node watcher, we'll check for node update
event and see if there are `out-of-service` taints are applied
and `ROOK_WATCH_FOR_NODE_FAILURE` is enabled in rook-ceph-operator-configmap,
if then we'll create the networkFence cr and delete the cr if nodes come back.
And, added the unit test too.
Signed-off-by: subhamkrai <srai@redhat.com>
The controller runtime logger has not been set, which means we are missing
information that could be useful in troubleshooting the controllers.
Now the logger is enabled to report this logging in the operator
log.
Signed-off-by: travisn <tnielsen@redhat.com>
Allow the validation tool to read test config from a yaml file. To help
users, also allow outputting a config file with default values and
comments instructing how to use the config file.
This work is in anticipation of adding more advanced configuration
options that would be too cumbersome to set using cli flags.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Allow overriding the nginx server image used for the web server and clients from CLI with --nginx-image flag.
Set default image in flag as 'nginxinc/nginx-unprivileged:stable-alpine'
Signed-off-by: iPraveenParihar <praveenparihar68@gmail.com>
Before starting multus validation test clients, pull the client image to
all nodes. This will ensure that variations in client readiness timing
will not be affected by variations in the time nodes take to pull the
image. This is intended to reduce the number of false reports of flaky
multus networks.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Use contents of KUBECONFIG var for creating Kubernetes client interface
if its present. This allows running rook CLI commands locally for
quicker development iteration.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Even though this is validated in the validation process, validating in
the cli gives better user feedback and doesn't spam the user with
suggestions for the failure that aren't relevant.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Add a more involved multus validation test to the Rook binary. Because
this is intended to be end-user runnable, make sure operator-only
commands are hidden.
Build this into the rook binary instead of creating a separate binary
for ease, and because any binary built with the kube api becomes 40+
megabytes. We save quite a bit of space by including this in the Rook
binary, which is good for keeping container layers as small as possible.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>