Dashboard ac-user-create cmd was taking more time
then the usual ceph command to run,
So increased the timeout to run the cmd, and
now dashboard admin user is sucessfully created
Closes: https://github.com/rook/rook/issues/12113
Signed-off-by: parth-gr <paarora@redhat.com>
Run the ganesha-rados-grace command in a remote pod when multus
networking is enabled.
Signed-off-by: parth-gr <paarora@redhat.com>
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Sometimes Ceph uses a different standard output to return errors or
merges standard error to standard out. So let's allow some commands to
return both in the output.
Signed-off-by: Sébastien Han <seb@redhat.com>
Previously we were merging the stderr even if it was empty, leading to
unmarshall errors.
The error simulation was done here
https://play.golang.org/p/Sk2yw9GUWNu.
Signed-off-by: Sébastien Han <seb@redhat.com>
When proxying commands to the cmd-proxy container we don't need to build
the command line with the same flags as the operator. The cmd-proxy
container does not use any ceph config file and just relies on the
CEPH_ARGS environment variable in the container. So passing the same
args as the operator causes to fail since we don't have a ceph config
file in `/var/lib/rook/openshift-storage/openshift-storage.config` thus
the remote exec fails with:
```
global_init: unable to open config file from search list ...
```
Signed-off-by: Sébastien Han <seb@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
We now pass the networking spec to the clusterInfo so that the executor
can make the right decision on how to execute a command.
This is a small change that allows us to remove the CephBlockPool CR
since it's checking for rbd images. On a Multus deployment, the operator
does not have the network annotations, thus has no access to the OSD
network then rbd commands are hanging forever.
Signed-off-by: Sébastien Han <seb@redhat.com>
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.
So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.
Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.
It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.
Signed-off-by: Sébastien Han <seb@redhat.com>
The integration tests have long been painful to maintain with
settings in various places and copied to multiple types,
inconsistent variable names, and otherwise difficult to maintain
code. Now the settings for a test suite are all in one place and
they remain in the same settings type throughout the test.
The multi-cluster suite is also refactored to use the same install
and uninstall helpers as the other suites.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When creating object store, `radosgw-admin realm get ..` command is stuck forever when required number of OSDs are not available. Because of this the uninstall of object store is also stuck. User has to manually remove the finalizer to delete the object store.This PR uses `ExecuteCommandWithTimeout` for running `radosgw-admin` command. Timeout during installation will be reconciled. Cleanup will be treated as best effort. Any errors during uninstalling of single site object store will only be logged.
Signed-off-by: Santosh Pillai <sapillai@redhat.com>
To provide more context for executing commands in a ceph cluster,
the full clusterInfo is now passed to the ceph execution commands.
All information about the cluster will now be available throughout
all the areas of the operator. The namespace, ceph credentials,
mon endpoints, and other info is a core part of that cluster info.
Arguments passed through the controllers are also simplified for
mons, mgr, osds, and other daemons where the parameters had
become too complex.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Currently, we need to configure the Ceph external Admin keyi
in the Rook deployment to be able to connect to an external Ceph cluster.
If we wanted to run in a multi-tenant fashion
were several K8s/Rook clusters wanted to connect to the same external ceph cluster,
each K8s deployment would have the access to the External Ceph Admin key
and could potentially access or delete the Data
from pools that belong to other k8s/Rook Clusters.
Now the admin key is optional but the helper script create-external-cluster-resources.sh
will help create the necessary keys/users to connect to that cluster.
Closes: https://github.com/rook/rook/issues/4917 and https://github.com/rook/rook/pull/5227
Signed-off-by: Sébastien Han <seb@redhat.com>
The PG count on metadata pools should default to rgw_rados_pool_pg_num_min
instead of the more general default pg count. This means rgw pools
will default to 8 PGs instead of 32 PGs, which means a lot more pools
can be created before hitting the default PG limit.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The ceph commands are now only written to the log in debug mode.
For commands that change the system state we now ensure that
a useful log entry is written. If all the details of the ceph
commands are needed, debug logging should still be enabled.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Allow users to specify overrides in the CephCluster CRD. The override
ConfigMap still exists for emergency situations and is mounted into
daemon pods directly instead of being merged into a Rook-created config
file.
Rook Ceph no longer uses a config file for managing daemons with the
exception of the OSDs which still generate a config in an init
container.
Signed-off-by: Blaine Gardner <blaine.gardner@suse.com>
this adds a structure to hold various settings related to executing ceph
cli commands, and introduces a common interface for configuring and
running such commands.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
The CephCluster resource will now expose the health of the ceph cluster
so admins can query the status with k8s api or kubectl instead of
needing to run the rook toolbox. The operator will query the
ceph status periodically and update the status in the custom resource.
Signed-off-by: travisn <tnielsen@redhat.com>
The Ceph CLI has the flag option '--connect-timeout' which controls the
timeout of the command when reaching out to a monitor. If the connection
cannot be established within 15 seconds, the command will fail. The
default value in Ceph is relatively long (300sec IIRC). This is a
problem when monitors have trouble reaching out quorum, since this means
the command will only fail after 300 sec.
Now we fail faster.
Resolves: https://github.com/rook/rook/issues/2574
Signed-off-by: Sébastien Han <seb@redhat.com>