Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.
So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.
Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.
It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.
Signed-off-by: Sébastien Han <seb@redhat.com>
In e8f9cfcb71, the logging was replaced by
Debug which is probably a mistake given the intent of the commit to
enable debug logging on the prepare job.
However, this code is only triggered when running a ceph-osd with the
rook binary, so we must run Info and run Debug since Debug is not
activated.
Signed-off-by: Sébastien Han <seb@redhat.com>
Work around issue https://github.com/rook/rook/issues/7573
and make sure integration tests check for regressions.
Eventually we should use the RADOS Gateway admin REST API, but for now
we need to work around an issue where the built version of
'radosgw-admin' has incompatibilities with the RADOS Gateway version
running in the Ceph cluster.
Of note, Rook built on the Ceph Pacific image will not support
some 'radosgw-admin' commands to Ceph Nautilus (v14) or Octopus (v15)
clusters.
Further complicating matters, the flag used for the workaround changes
between Ceph v16.2.0 and v16.2.1 (both Pacific).
This bug needs to be treated a little differently than most of the ways
Rook handles different commands for different Ceph versions because this
is based on the Ceph version that is installed in the container with the
Rook operator primarily and not the version of Ceph running in the
cluster.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Somehow the combined output does not contain the stderr in the errorn
only in the buffer.
For example:
err = "error exit 1"
buf = "Error EBUSY: not enough monitors would be available () after stopping mons [a]"
So we now embed the buf in the error so that the caller does not need to
print the buffer as well as the error.
Also, the caller can more easily decide to ignore the error by
instropecting the string.
Signed-off-by: Sébastien Han <seb@redhat.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
We don't always get an error of the type "*exec.ExitError" so we must
validate the type before printing it otherwise the interface conversion
will fail.
Signed-off-by: Sébastien Han <seb@redhat.com>
In order to properly debug errors, we need to merge stderr inside the
`err` reported so that we don't only see the stdout.
We were doing this when executing without file output, doing the same
for the output file.
Signed-off-by: Sébastien Han <seb@redhat.com>
using defer for closing file which are open
for writing is not safe. so closing file again
following below steps:
1. open files
2. defer file.close()
3. write
4.file.close()
these will make sure files are closed.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
this commit suppress the gosec errors for
g204: Audit use of command execution.
g304: File path provided as taint input.
g101: Look for hard coded credentials.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
To ensure a file handle is closed, we defer the close command
so it is guaranteed to run when the method returns. The closing
of the file handle is not going to fail in our usage since we aren't
using the SetDeadline on the files that would cancel a request
and return an error.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The methods and arguments to the exec methods are not all used anymore.
This cleans up the methods to only what is necessary to improve
the readability and maintainability.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:
* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion
Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
Signed-off-by: Guy Margalit <guymguym@gmail.com>
Co-Authored-By: Sébastien Han <seb@redhat.com>
Co-Authored-By: Travis Nielsen <tnielsen@redhat.com>
This change is meant to allow running operators locally on a developer machine.
The idea is to allow faster development cycles by reducing the time and complexity of building -> deploying -> debugging on cluster.
For operators that rely only on kubernetes API this works easily - see cockroachdb and minio examples in development-flow doc.
The change includes:
- rook.NewContext() - Refactored to remove repeating initialization code that was copy-pasted in most of the operators in order to create the clusterd.Context and the Clientsets. Also it detects the mode of working in-cluster vs external and sets up the external mode with standard user config (~/.kube/config) and a job executor.
- rook.GetOperatorImage() - Refactor this repeating code in many operators to detect the operator pod image. Also added a global flag --operator-image that developers can use to override this when running locally.
- rook.TerminateOnError() - Added a convenient function.
the existing exec interface with timeout is effectively the same as
ExecuteCommandWithOutput plus a timeout. this patch adds a variant of
ExecuteCommandWithOutputFile that uses a timeout.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
The io.MultiReader(r1, r2) reads r1 until EOF before moving on to r2.
When used for stdout/stderr stderr will not be written to the log until
stdout reaches EOF. This patch reads stderr in a go routine so that the
two streams can be interleaved properly.
Fixes: #2479
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
Previous code tried to interpret the string output as a `fmt.Sprintf`
string formatter, mangling `%` into `%!(MISSING)`, etc.
Thanks to go, I don't belive these are exploitable (unlike the similar
error in C).
Since this seemed to be a common error in the codebase, I did a quick
audit by visually inspecting the results of `git grep 'f([^"]'`. I
don't have a good suggestion for automated tests to prevent this in
future :(
Example error (look for `(MISSING)`):
```
E0927 05:31:07.618429 11227 driver-call.go:237] Failed to unmarshal output for command: unmount, output: "2018-09-27 05:31:07.191711 I | exec: Running command: df --type ceph /var/lib/kubelet/pods/95461479-c216-11e8-bcf0-02030782ac80/volumes/ceph.rook.io~rook/oe-scratch\n2018-09-27 05:46:43.808596 I | Filesystem 1K-blocks Used Available Use%!M(MISSING)ounted on\n2018-09-27 05:46:43.808659 I | 10.107.25.147:6790,10.109.173.79:6790,10.104.85.255:6790:/ 151678976 49410048 102268928 33%!/(MISSING)var/lib/kubelet/pods/95461479-c216-11e8-bcf0-02030782ac80/volumes/ceph.rook.io~rook/oe-scratch\n{\"status\":\"Success\"}\n", error: invalid character '-' after top-level value
```
Signed-off-by: Angus Lees <gus@inodes.org>
- Operator is deployed under namespace rook-system
- Operator deploys rook-agent on all nodes as daemonset also on
rook-system
- Rook-agent installs flexvolume driver on hosts
- Rook-agent listens on unix socket for driver requests
- Rook-agent perform attach/detach on its node
- Rook-agent creates/delete CRD volumeattach objects
- Added fencing to support ROX and RWO
- Added unit and integration tests
- Updated examples and docks
fixes#432
RBD image support fixes, ensure /etc/ceph is created in toolbox image, return output string in exec even for errors because the ceph tools return useful information there upon error