Commit Graph
44 Commits
Author SHA1 Message Date
Satoru Takeuchi 30e4fbb01f ceph: make the timeout of ceph commands cofigurable
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.

Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
2021-07-27 12:42:16 +00:00
Sébastien Han 6d77a9976c ceph: remove unnecessary exec helpers
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.

Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-23 09:16:33 +02:00
Sébastien Han bd58790c31 ceph: proxy ceph commands when multus is configured
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.

So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.

Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.

It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-07 19:08:32 +02:00
Sébastien Han 8c589a3785 ceph: silence harmless errors
Let's not look at misleading errors if the operator is still
initializing.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-07-01 15:37:08 +02:00
Sébastien Han da56354aa3 ceph: run CLI output with Info logging
In e8f9cfcb71, the logging was replaced by
Debug which is probably a mistake given the intent of the commit to
enable debug logging on the prepare job.
However, this code is only triggered when running a ceph-osd with the
rook binary, so we must run Info and run Debug since Debug is not
activated.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-06-22 16:43:47 +02:00
Blaine Gardner d78b2a41d6 ceph: work around radosgw-admin fifo file io error
Work around issue https://github.com/rook/rook/issues/7573
and make sure integration tests check for regressions.

Eventually we should use the RADOS Gateway admin REST API, but for now
we need to work around an issue where the built version of
'radosgw-admin' has incompatibilities with the RADOS Gateway version
running in the Ceph cluster.

Of note, Rook built on the Ceph Pacific image will not support
some 'radosgw-admin' commands to Ceph Nautilus (v14) or Octopus (v15)
clusters.

Further complicating matters, the flag used for the workaround changes
between Ceph v16.2.0 and v16.2.1 (both Pacific).

This bug needs to be treated a little differently than most of the ways
Rook handles different commands for different Ceph versions because this
is based on the Ceph version that is installed in the container with the
Rook operator primarily and not the version of Ceph running in the
cluster.

Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
2021-04-13 12:08:02 -06:00
Sébastien Han 2d75135579 ceph: embed the buffer in the error
Somehow the combined output does not contain the stderr in the errorn
only in the buffer.
For example:

err = "error exit 1"
buf = "Error EBUSY: not enough monitors would be available () after stopping mons [a]"

So we now embed the buf in the error so that the caller does not need to
print the buffer as well as the error.
Also, the caller can more easily decide to ignore the error by
instropecting the string.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-02-03 17:28:47 +01:00
Sébastien Han b189e91bfe ceph: use errors pkg instead of fmt
This one was a leftover.

Signed-off-by: Sébastien Han <seb@redhat.com>
2021-02-03 17:28:47 +01:00
Mateusz Gozdek 8ba3762fa4 docs: fix bunch of typos
Found by running the following command:

codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H

Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
2020-11-06 10:01:04 +01:00
Sébastien Han 1168dbc6ca ceph: assert error type before using it
We don't always get an error of the type "*exec.ExitError" so we must
validate the type before printing it otherwise the interface conversion
will fail.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-28 17:08:38 +01:00
Sébastien Han 9b980d136f ceph: log stderr in error when exec with output file
In order to properly debug errors, we need to merge stderr inside the
`err` reported so that we don't only see the stdout.
We were doing this when executing without file output, doing the same
for the output file.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-10-07 14:47:09 +02:00
subhamkrai c938849cf8 ceph: closing file which are open for writing
using defer for closing file which are open
for writing is not safe. so closing file again
following  below steps:
1. open files
2. defer file.close()
3. write
4.file.close()

these will make sure files are closed.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-30 22:37:56 +05:30
subhamkrai 0279025e9e ceph: handling all the gosec errors
a few of the gosec errors were left. so
this commit will resolve all the errors.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-30 08:19:55 +05:30
subhamkrai bcd7faed4e ceph: suppress gosec errors for g204, g304, g101
this commit suppress the gosec errors for

g204: Audit use of command execution.
g304: File path provided as taint input.
g101: Look for hard coded credentials.

Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
2020-07-28 11:10:48 +05:30
Travis Nielsen cf53467380 core: suppress gosec errors for closing files
To ensure a file handle is closed, we defer the close command
so it is guaranteed to run when the method returns. The closing
of the file handle is not going to fail in our usage since we aren't
using the SetDeadline on the files that would cancel a request
and return an error.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-07-24 12:31:13 -06:00
Sébastien Han fa32f5323d ceph: fix exec output
The previous code was overriding the content of `out`, now we just
return after the error.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-25 17:05:07 +02:00
Sébastien Han 5f74e493ef ceph: small user delete refactor
Do not return error code, interpret it directly, put the error as part
of the output on failures.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-18 18:49:03 +02:00
Sébastien Han 84d1e28c99 ceph: add external support for objectstoreuser
Now, the object store user is capable of creating s3 users on an
external Ceph cluster.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-06-18 16:34:01 +02:00
Sébastien Han e96dc646a2 rook: add ExecuteCommandWithEnv
We can now execute commands and pass env variables to the executor.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-20 17:37:58 +01:00
Travis Nielsen e8f9cfcb71 exec: always write commands to debug log
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:54 -06:00
Travis Nielsen f2ecaa2bda exec: remove the unused actionName param
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Travis Nielsen 84b8cdcf75 exec: simplify the exec package from unused methods and logging
The methods and arguments to the exec methods are not all used anymore.
This cleans up the methods to only what is necessary to improve
the readability and maintainability.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2020-03-19 07:49:53 -06:00
Sébastien Han a3068dee0b ceph: separate controller for CephBlockPool CRD
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:

* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion

Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-02 17:31:03 +01:00
Sébastien Han 65e1c054a3 ceph: only print command when running debug mode
We don't need commands e run under the hood. Enable debug logs to see
that.

Signed-off-by: Sébastien Han <seb@redhat.com>
2020-03-02 09:29:48 +01:00
Guangming Wang e5e0369e37 cleanup: cleanup codebase
exec.go: use String method of bytes itself.

Signed-off-by: Guangming Wang <guangming.wang@daocloud.io>
2019-11-07 19:45:32 +08:00
Guy Margalit 2ce2382676 Refactor operator context init and allow to run local
Signed-off-by: Guy Margalit <guymguym@gmail.com>
Co-Authored-By: Sébastien Han <seb@redhat.com>
Co-Authored-By: Travis Nielsen <tnielsen@redhat.com>

This change is meant to allow running operators locally on a developer machine.
The idea is to allow faster development cycles by reducing the time and complexity of building -> deploying -> debugging on cluster.

For operators that rely only on kubernetes API this works easily - see cockroachdb and minio examples in development-flow doc.

The change includes:

- rook.NewContext() - Refactored to remove repeating initialization code that was copy-pasted in most of the operators in order to create the clusterd.Context and the Clientsets. Also it detects the mode of working in-cluster vs external and sets up the external mode with standard user config (~/.kube/config) and a job executor.
- rook.GetOperatorImage() - Refactor this repeating code in many operators to detect the operator pod image. Also added a global flag --operator-image that developers can use to override this when running locally.
- rook.TerminateOnError() - Added a convenient function.
2019-06-19 20:47:55 +03:00
Noah Watkins b29ba26b12 exec: add timeout variant of exec with output file
the existing exec interface with timeout is effectively the same as
ExecuteCommandWithOutput plus a timeout. this patch adds a variant of
ExecuteCommandWithOutputFile that uses a timeout.

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-05-06 10:10:05 -07:00
Noah Watkins fbb56c41ed exec: interleave stdout/stderr logging
The io.MultiReader(r1, r2) reads r1 until EOF before moving on to r2.
When used for stdout/stderr stderr will not be written to the log until
stdout reaches EOF. This patch reads stderr in a go routine so that the
two streams can be interleaved properly.

Fixes: #2479

Signed-off-by: Noah Watkins <noahwatkins@gmail.com>
2019-01-24 13:10:46 -08:00
Angus Lees 9c95ed5f6f Don't misinterpret command output as string format
Previous code tried to interpret the string output as a `fmt.Sprintf`
string formatter, mangling `%` into `%!(MISSING)`, etc.

Thanks to go, I don't belive these are exploitable (unlike the similar
error in C).

Since this seemed to be a common error in the codebase, I did a quick
audit by visually inspecting the results of `git grep 'f([^"]'`.  I
don't have a good suggestion for automated tests to prevent this in
future :(

Example error (look for `(MISSING)`):
```
E0927 05:31:07.618429   11227 driver-call.go:237] Failed to unmarshal output for command: unmount, output: "2018-09-27 05:31:07.191711 I | exec: Running command: df --type ceph /var/lib/kubelet/pods/95461479-c216-11e8-bcf0-02030782ac80/volumes/ceph.rook.io~rook/oe-scratch\n2018-09-27 05:46:43.808596 I | Filesystem                                                 1K-blocks     Used Available Use%!M(MISSING)ounted on\n2018-09-27 05:46:43.808659 I | 10.107.25.147:6790,10.109.173.79:6790,10.104.85.255:6790:/ 151678976 49410048 102268928  33%!/(MISSING)var/lib/kubelet/pods/95461479-c216-11e8-bcf0-02030782ac80/volumes/ceph.rook.io~rook/oe-scratch\n{\"status\":\"Success\"}\n", error: invalid character '-' after top-level value
```

Signed-off-by: Angus Lees <gus@inodes.org>
2018-10-15 15:59:13 +11:00
Jared Watts 6a2e20aaa8 include stderr output in CommandError.Error() output
Signed-off-by: Jared Watts <jbw976@gmail.com>
2018-03-28 13:26:00 -07:00
Steve Leon b4279e06fb Rook plugin for Kubernetes implemented as flexvolume
- Operator is deployed under namespace rook-system
- Operator deploys rook-agent on all nodes as daemonset also on
  rook-system
- Rook-agent installs flexvolume driver on hosts
- Rook-agent listens on unix socket for driver requests
- Rook-agent perform attach/detach on its node
- Rook-agent creates/delete CRD volumeattach objects
- Added fencing to support ROX and RWO
- Added unit and integration tests
- Updated examples and docks

fixes #432
2017-10-09 09:14:08 -07:00
Travis Nielsen ec181c0390 logs: stop writing the frequent health check 2017-08-07 10:58:52 -07:00
Travis Nielsen cb5741989b log all output of commands executed with an output file 2017-06-21 11:29:14 -07:00
Jared Watts 585501b925 use an output file when invoking ceph commands to isolate the payload from any logging 2017-06-20 20:57:46 -07:00
Travis Nielsen 0a035ed402 support for RBD images, all api tests passing
RBD image support fixes, ensure /etc/ceph is created in toolbox image, return output string in exec even for errors because the ceph tools return useful information there upon error
2017-06-20 20:57:46 -07:00
Travis Nielsen 58aa72a254 remove dependency on bash 2017-05-05 12:25:31 -07:00
Jared Watts e5cca80654 Operator and cephmgr support for specifying storage resources and configuration 2017-04-04 15:32:49 -07:00
Michael Goff 4870f951f0 cli and rgw: Added support for list, create, update, get, and delete users as well as listing buckets. Updated connection info to be for a user. 2017-02-08 09:50:19 -08:00
Travis Nielsen 1d2cca0a26 check if process is being monitored before starting a new process 2016-12-16 09:11:26 -08:00
Travis Nielsen cd21a158d8 log child process output with specific log id 2016-11-18 22:45:21 -08:00
Jared Watts 0e00f57758 Migrate logging package from built-in to github.com/coreos/capnslog, allow user specified logging level, tune logging levels for existing log statements 2016-11-16 14:19:42 -08:00
Bassam Tabbara b3c4f18c89 initial docs and licensing 2016-11-07 17:13:39 -08:00
Travis Nielsen 632f418cea execute all child processes with logging and same executor 2016-10-31 13:07:11 -07:00
Jared Watts 4e6b5424b2 Move Executor functionality to its own package 2016-10-12 09:07:47 -07:00