Two call sites build a shell command by interpolating values that were
never shell-quoted, so a value holding a metacharacter is executed
rather than passed along: RadosNamespaceHasObjects joins the finalized
rados argument vector with spaces, and CopyLocalFileToContainer
interpolates the destination path into a redirect.
Neither has misbehaved so far because rook composes both values itself,
not because the inputs are constrained to shell-safe characters.
CopyLocalFileToContainer receives a temp-file path. The rados probe
receives a "<metadataPoolName>:<zone><suffix>" string assembled by
adjustZoneDefaultPools, whose pool half comes from the default pool
placement's metadataPoolName or from spec.sharedPools.metadataPoolName,
neither of which carries pattern validation, or from the zone JSON of a
zone rook adopted rather than created.
Quote both with FormatCommand.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
A command logged by the exec package could not be copied out of the log
and re-run. Creating the RGW admin ops user, for example, logged an
unquoted "--display-name RGW Admin Ops User" and an unquoted "--caps
accounts=*;buckets=*;users=*;usage=read;metadata=read;zone=read", which
a shell reads as four arguments and six commands respectively. Render
these lines with FormatCommand instead.
cmd-reporter also named the command twice, because exec.Cmd already
carries argv[0] in Args and the log line prefixed Path to the whole of
Args.
The startup line drops the quotes that wrapped the entire argument list,
now that each argument carries its own; both in-tree samples of that
line are updated to match.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Command arguments are passed to the kernel as a vector, so joining them
with plain spaces loses the boundaries between them. An argument holding
whitespace or a shell metacharacter reads back as several arguments, or
as several commands.
FormatCommand renders a command and its arguments as a single line a
POSIX shell reads back as the original vector. Quoting is minimal: an
argument is single-quoted only when it contains a character the shell
would not pass through literally, so the flag-heavy command lines rook
builds stay readable.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
executeCommandWithTimeout wires a single bytes.Buffer to both cmd.Stdout
and cmd.Stderr and joins the command in a goroutine via cmd.Wait(). On
the timeout kill path it called cmd.Process.Kill() and then read that
buffer without waiting for cmd.Wait() to return. Kill() only signals the
process; it does not wait for os/exec's output-copier goroutines (joined
only by cmd.Wait) to finish, so reading the buffer there races with those
writers, which the race detector flags.
Only read the buffer once the copier goroutines have drained, signaled by
cmd.Wait() on the done channel. Because a killed process can leave an
orphaned descendant holding the output pipe open -- a D-state
cryptsetup/dmsetup child, or the downstream of a "sh -c '... | head'"
pipeline -- cmd.Wait() may never return, so bound the drain with a short
grace period and give up on the captured output rather than blocking the
caller forever. This keeps the bounded return the timeout path is meant
to guarantee. On a failed Kill() the process may still be writing, so
return without reading the buffer.
Add a regression test that drives the kill path with a child that ignores
SIGINT while writing to stdout; it fails under -race before this change.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Creating the rgw admin user is taking on the order of two
minutes instead of the expected seconds, thus the command
was timing out and cause the object store user creation
to fail. The increased timeout allows the admin user creation
to complete, and other users to be created as well.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Update the unit test for ExecuteCommandWithTimeout(). This unit failed a
CI run. This seems to have been a race condition where the previous
`cat` command was returning quickly enough that it was available to the
`select` statment at the same time as the timeout.
Resolve this issue by replacing the `cat` command with `sleep`, which
will be guaranteed to take longer to return than the timeout, ensuring
this race doesn't happen.
While working on the test, I also notice that the 30 second timeouts for
other tests would be too long in the event the commands hang for any
reason, so lower these to help ensure the CI won't time out due to
unforeseen events.
Also add a missing unit test to ensure error returns (via `false`
command) without timeouts are properly returned.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.
Additionally, it silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
and string format errors found by golangci-lint, while at it.
Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
Due to multi-site Ceph object resources (realms, zonegroups, zones)
reconciliation CLI commands being proxied through CephCluster's Mgr pods
in Multus-based scenarios, return codes were being additionally wrapped
by K8S' command streams and weren't adequately handled in-code, causing
reconciliation of CephObject(Realm|ZoneGroup|Zone) resources to fail, as
an exit code of 0, instead of the expected syscall.ENOENT (2) was returned.
Fixes#12833
Signed-off-by: zer0def <zv0uzqbqnncivw0afjsx79jmd19shpps3r5f2vjc6yv@definedaszero.xyz>
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
This commit implements this corner case during osd-prepare job.
```
The encrypted block is not opened, this is an extreme corner case
The OSD deployment has been removed manually AND the node rebooted
So we need to re-open the block to re-hydrate the OSDInfo.
Handling this case would mean, writing the encryption key on a
temporary file, then call luksOpen to open the encrypted block and
then call ceph-volume to list against the opened encrypted block.
We don't implement this, yet and return an error.
```
When underlying PVC for osd are CSI provisioned, the encrypted device
is closed when PVC is unmounted due to osd pod being deleted.
Therefore, this may occur more frequently and needs to be handled.
This commit implements the fix for the same.
Signed-off-by: Rakshith R <rar@redhat.com>
The tool gofmt in go 1.19 requires certain formatting
in the comments section for better rendering.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
In volume.go, remove if condition in `callCephVolume`
method where we always pass false.
In exec.go, removing `ExecuteCommandWithFile*` methods
which was not used anywhere.
Signed-off-by: subhamkrai <srai@redhat.com>
Rook should update the RGW object store's period if the period doesn't
yet exist. This protects us from the case where the
'radosgw-admin period update --commit` command fails and the
CephObjectStore controller reconciles again.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
In ExtractExitCode, there is one error type that can be valid as a
pointer or not-as-a-pointer. Add a case to the type check for the
non-pointer condition.
Resolves#8280
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Sometimes the default 15s is not enough for timeout of ceph commands. For examples,
I encountered that `radosgw-admin` command took dozens of seconds under heavy load.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
Both `ExecuteCommandWithOutputFileTimeout()` and
`ExecuteCommandWithOutputFile()` generate unnecessary system calls by
creating/reading/removing files where the stream output of the command
can simply be used. So sticking with `ExecuteCommandWithOutput()` and
`ExecuteCommandWithCombinedOutput()` for reading outputs is sufficient.
Closes: https://github.com/rook/rook/issues/8343
Signed-off-by: Sébastien Han <seb@redhat.com>
When the CephCluster is configured with Multus and multiple networks are
used to deploy Ceph some commands are failing to be executed from the
Operator. These commands, in particular, `radosgw-admin` ones need access
to the "ceph public network" to talk to OSDs. Unfortunately, the
Rook-Ceph Operator does not have the network annotations and thus
doesn't have the networks available and cannot reach OSDs. So the commands end
up hanging and eventually time out.
Applying the annotations to the Operator pod is possible but will result
in restarting the operator too and this should be avoided at all costs.
Also, applying the annotations beforehand is not possible since the
Multus declaration is in the CephCluster specification. So we would have
no idea what to do.
So the current approach runs a new sidecar container in the mgr pod to
act as a proxy for "some" ceph commands, only the `radosgw-admin` ones
for multi-site setup. This is a small container with admin access
running idle waiting for commands to be executed. In a sense, it is
similar to the toolbox but we didn't want to clearly expose it, so
running as a sidecar is quite nice.
Proxying command is obviously not always recommended since we add an
extra hop in the network path. Now each request has to go from the
operator pod to the API server to the remote pod to Ceph. Previously,
the command only goes from the operator to Ceph.
It's worth noting that external mode is not impacted since no rgw pod
is configured. This scenario is flexible and allows us to scale
pretty well since any CephCluster with Multus will see its mgr sidecar
deployed and can then talk to Ceph. We are not limited.
Signed-off-by: Sébastien Han <seb@redhat.com>
In e8f9cfcb71, the logging was replaced by
Debug which is probably a mistake given the intent of the commit to
enable debug logging on the prepare job.
However, this code is only triggered when running a ceph-osd with the
rook binary, so we must run Info and run Debug since Debug is not
activated.
Signed-off-by: Sébastien Han <seb@redhat.com>
Work around issue https://github.com/rook/rook/issues/7573
and make sure integration tests check for regressions.
Eventually we should use the RADOS Gateway admin REST API, but for now
we need to work around an issue where the built version of
'radosgw-admin' has incompatibilities with the RADOS Gateway version
running in the Ceph cluster.
Of note, Rook built on the Ceph Pacific image will not support
some 'radosgw-admin' commands to Ceph Nautilus (v14) or Octopus (v15)
clusters.
Further complicating matters, the flag used for the workaround changes
between Ceph v16.2.0 and v16.2.1 (both Pacific).
This bug needs to be treated a little differently than most of the ways
Rook handles different commands for different Ceph versions because this
is based on the Ceph version that is installed in the container with the
Rook operator primarily and not the version of Ceph running in the
cluster.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Somehow the combined output does not contain the stderr in the errorn
only in the buffer.
For example:
err = "error exit 1"
buf = "Error EBUSY: not enough monitors would be available () after stopping mons [a]"
So we now embed the buf in the error so that the caller does not need to
print the buffer as well as the error.
Also, the caller can more easily decide to ignore the error by
instropecting the string.
Signed-off-by: Sébastien Han <seb@redhat.com>
Found by running the following command:
codespell -S .git,*.png,*.jpg -L \
aks,keyserver,atleast,dne,ser,ist,files\',ba,dum,iam,te -f -H
Signed-off-by: Mateusz Gozdek <mgozdekof@gmail.com>
We don't always get an error of the type "*exec.ExitError" so we must
validate the type before printing it otherwise the interface conversion
will fail.
Signed-off-by: Sébastien Han <seb@redhat.com>
In order to properly debug errors, we need to merge stderr inside the
`err` reported so that we don't only see the stdout.
We were doing this when executing without file output, doing the same
for the output file.
Signed-off-by: Sébastien Han <seb@redhat.com>
using defer for closing file which are open
for writing is not safe. so closing file again
following below steps:
1. open files
2. defer file.close()
3. write
4.file.close()
these will make sure files are closed.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
this commit suppress the gosec errors for
g204: Audit use of command execution.
g304: File path provided as taint input.
g101: Look for hard coded credentials.
Signed-off-by: subhamkrai <subhamkumarrai03@gmail.com>
To ensure a file handle is closed, we defer the close command
so it is guaranteed to run when the method returns. The closing
of the file handle is not going to fail in our usage since we aren't
using the SetDeadline on the files that would cancel a request
and return an error.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The helpers for executing a process have long required an actionName
param which is not being used. Now we remove the old param
while also cleaning up various other usages of the exec
package.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
The methods and arguments to the exec methods are not all used anymore.
This cleans up the methods to only what is necessary to improve
the readability and maintainability.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Now, the CephBlockPool CRD is managed with the controller-runtime.
So the watcher is outside of the main controller reconciliation loop of
CephCluster which brings numerous benefit such as:
* having its own reconciliation loop
* won't block anything from the main CephCluster controller loop
* fast than waiting for CephCluster loop to completion
Partially close: https://github.com/rook/rook/issues/1981
Signed-off-by: Sébastien Han <seb@redhat.com>
Signed-off-by: Guy Margalit <guymguym@gmail.com>
Co-Authored-By: Sébastien Han <seb@redhat.com>
Co-Authored-By: Travis Nielsen <tnielsen@redhat.com>
This change is meant to allow running operators locally on a developer machine.
The idea is to allow faster development cycles by reducing the time and complexity of building -> deploying -> debugging on cluster.
For operators that rely only on kubernetes API this works easily - see cockroachdb and minio examples in development-flow doc.
The change includes:
- rook.NewContext() - Refactored to remove repeating initialization code that was copy-pasted in most of the operators in order to create the clusterd.Context and the Clientsets. Also it detects the mode of working in-cluster vs external and sets up the external mode with standard user config (~/.kube/config) and a job executor.
- rook.GetOperatorImage() - Refactor this repeating code in many operators to detect the operator pod image. Also added a global flag --operator-image that developers can use to override this when running locally.
- rook.TerminateOnError() - Added a convenient function.
the existing exec interface with timeout is effectively the same as
ExecuteCommandWithOutput plus a timeout. this patch adds a variant of
ExecuteCommandWithOutputFile that uses a timeout.
Signed-off-by: Noah Watkins <noahwatkins@gmail.com>