Adding callback function in the osd removal method
as in downstream there is requirement of adding extra
check before proceeding with osd removal.
Signed-off-by: subhamkrai <srai@redhat.com>
Add the ability to specify node profiles in the multus validation test.
This addresses a few points of early feedback on the validation tool.
Statements below critique the tool's behavior before this patch.
1. The tool assumes all daemons are on public and cluster network, which
means users who have a significantly smaller cluster net (a
design choice) cannot run a single test to determine if Rook is
likely to install correctly.
2. The tool does not have placement options to select only a subset of
Kubernetes nodes to run validation on.
3. Users of multus seem to have a dedicated pool of storage nodes more
often than the average Rook install. This makes sense for security-
and perforance-minded users. The tool cannot run a single test to
verify storage-only and general-workload nodes at one time.
These points are addressed by allowing users to specify configurations
for different "NodeTypes."
Each NodeType config has options for selecting the number of OSDs as
well as the number of other (non-OSD) Ceph daemons. This limits the
unnecessary exhaustion of cluster network addresses from critique 1.
Each NodeType config has its own placement (critique 2).
Users can define as many NodeTypes as needed to test the network for
their planned CephCluster. Specifically, this allows the tool to test
storage-only nodes and generalized-workload nodes at the same time. An
arbitrary number of NodeTypes are allowed to support even more highly
specialized cluster setups, such as multiple tiers of storage nodes
where some storage-only nodes may run more OSDs than others.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
If osd store is updated in the ceph cluster, then
delete OSDs one by one, cleanup disks and provision a new OSD on
the same disk
Signed-off-by: sp98 <sapillai@redhat.com>
The function ParseMonEndpoints existed twice, once in the mon package
and once in the controller package. It has now been deleted from the
mon package. Usages in the mon package now reference the function
in the controller package.
Signed-off-by: Henry Davies <henrydavies@hotmail.co.uk>
The previous client loader routine assumed KUBECONFIG would be set in
CLI environments. Instead, now take this approach:
1. If KUBECONFIG is set, that is the de-facto override that informs the
tool it is being run in a CLI environment.
2. Otherwise, try creating a client from the default kube config file.
3. If that fails, assume the tool is running in a Kubernetes Pod.
This will continue supporting dev/test environments so that building the
rook container is not necessary for development. It also allows support
for highly opinionated environments (like deployed by OpenShift CSVs)
where it's not possible to install the Rook operator without also
deploying a CephCluster.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
in the existing node watcher, we'll check for node update
event and see if there are `out-of-service` taints are applied
and `ROOK_WATCH_FOR_NODE_FAILURE` is enabled in rook-ceph-operator-configmap,
if then we'll create the networkFence cr and delete the cr if nodes come back.
And, added the unit test too.
Signed-off-by: subhamkrai <srai@redhat.com>
The controller runtime logger has not been set, which means we are missing
information that could be useful in troubleshooting the controllers.
Now the logger is enabled to report this logging in the operator
log.
Signed-off-by: travisn <tnielsen@redhat.com>
Allow the validation tool to read test config from a yaml file. To help
users, also allow outputting a config file with default values and
comments instructing how to use the config file.
This work is in anticipation of adding more advanced configuration
options that would be too cumbersome to set using cli flags.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Allow overriding the nginx server image used for the web server and clients from CLI with --nginx-image flag.
Set default image in flag as 'nginxinc/nginx-unprivileged:stable-alpine'
Signed-off-by: iPraveenParihar <praveenparihar68@gmail.com>
Before starting multus validation test clients, pull the client image to
all nodes. This will ensure that variations in client readiness timing
will not be affected by variations in the time nodes take to pull the
image. This is intended to reduce the number of false reports of flaky
multus networks.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Use contents of KUBECONFIG var for creating Kubernetes client interface
if its present. This allows running rook CLI commands locally for
quicker development iteration.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Even though this is validated in the validation process, validating in
the cli gives better user feedback and doesn't spam the user with
suggestions for the failure that aren't relevant.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Add a more involved multus validation test to the Rook binary. Because
this is intended to be end-user runnable, make sure operator-only
commands are hidden.
Build this into the rook binary instead of creating a separate binary
for ease, and because any binary built with the kube api becomes 40+
megabytes. We save quite a bit of space by including this in the Rook
binary, which is good for keeping container layers as small as possible.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
this feature is to allow the use of filters using the deviceFilter flag
by skipping devices that do not match the filter
Closes: #10340
Signed-off-by: Javier <sjavierlopez@gmail.com>
This commit adds functionality to be able to rotate
key encryption key of encrypted PVC backed OSDs.
Necessary changes such as adding update functionality
to kms and rbac changes are made as well.
Signed-off-by: Rakshith R <rar@redhat.com>
With the new mgr HA implementation (based on readiness probe) we
don't need the mgr sidecar (live-watch) anymore. This commit is
intended to remove all the related code.
Signed-off-by: Redouane Kachach <rkachach@redhat.com>
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
Environment variables are not recommended for secrets in pods since
they can be easily leaked if the environemnt variables are logged.
By mounting the mon secret as a file, the mgr and osd prepare pods
can read the mon secret from a file for better security.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The fsid, username, and user secrets were from legacy and no
longer needed on the osd daemon or mon pods. This removes the
unnecessary env vars from the pod specs.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The tool gofmt in go 1.19 requires certain formatting
in the comments section for better rendering.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The RGW support server side encryption with help of s3 protocol, till
now the `sse:kms` was support in which keys will be provided by the user
and but it will be saved in external management service like vault. Now
the support for `sse:s3` is added so the entire encryption key
management is performed by RGW itsels.
Signed-off-by: Jiffin Tony Thottan <thottanjiffin@gmail.com>
The cluster-wide encryption feature now has a new supported key
management system: IBM Key protect. More information about the backend
can be found in Rook documentation under the
Today's implementation stores OSD encryption keys has "Standard" keys.
Signed-off-by: Sébastien Han <seb@redhat.com>
Previously we were using the namespace, it's a mistake, even though most
clusters use the same name as the namespace.
Let's be precise and use the cluster name when looking for it.
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit adds context parameter to k8sutil kvstore functions. By
this, we can handle cancellation during API call of kvstore resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
If multiple removal jobs are fired in parallel, there is a risk of
losing data since we will forcefully remove the OSD. It's also simply
true if a single OSD is not safe to destroy, there is also a risk of
data loss.
So now, we check if the OSD is safe-to-destroy first and then proceed.
The code waits forever and retries every minute unless the
--force-osd-removal flag is passed.
Signed-off-by: Sébastien Han <seb@redhat.com>
The cluster info is important context for the cluster controller to
create the cluster, and all the fields must be properly set.
A test cluster name was being set temporarily, resulting in
mons incorrectly getting the wrong cluster CR name. There is no
known issue from the temporary value, it was just exposed by
https://github.com/rook/rook/pull/8678 setting the value to a label.
Now the functions are more clearly named so only unit and
integration tests should be using the test value for the cluster
name where it is not important.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
This commit adds context parameter to k8sutil node functions. By this,
we can handle cancellation during API call of node resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>
rook command doesn't interpret `logtostderr` option. It's OK to just
remove this option because `capnslog` outputs all logs to stdout
by default. It's better to keep `AddGoFlagSet()` call because
some libraries might define their own flags with Go's `flag` package.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
This commit adds context parameter to k8sutil pod functions. By this, we
can handle cancellation during API call of pod resource.
Signed-off-by: Yuichiro Ueno <y1r.ueno@gmail.com>