executeCommandWithTimeout wires a single bytes.Buffer to both cmd.Stdout
and cmd.Stderr and joins the command in a goroutine via cmd.Wait(). On
the timeout kill path it called cmd.Process.Kill() and then read that
buffer without waiting for cmd.Wait() to return. Kill() only signals the
process; it does not wait for os/exec's output-copier goroutines (joined
only by cmd.Wait) to finish, so reading the buffer there races with those
writers, which the race detector flags.
Only read the buffer once the copier goroutines have drained, signaled by
cmd.Wait() on the done channel. Because a killed process can leave an
orphaned descendant holding the output pipe open -- a D-state
cryptsetup/dmsetup child, or the downstream of a "sh -c '... | head'"
pipeline -- cmd.Wait() may never return, so bound the drain with a short
grace period and give up on the captured output rather than blocking the
caller forever. This keeps the bounded return the timeout path is meant
to guarantee. On a failed Kill() the process may still be writing, so
return without reading the buffer.
Add a regression test that drives the kill path with a child that ignores
SIGINT while writing to stdout; it fails under -race before this change.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Update the unit test for ExecuteCommandWithTimeout(). This unit failed a
CI run. This seems to have been a race condition where the previous
`cat` command was returning quickly enough that it was available to the
`select` statment at the same time as the timeout.
Resolve this issue by replacing the `cat` command with `sleep`, which
will be guaranteed to take longer to return than the timeout, ensuring
this race doesn't happen.
While working on the test, I also notice that the 30 second timeouts for
other tests would be too long in the event the commands hang for any
reason, so lower these to help ensure the CI won't time out due to
unforeseen events.
Also add a missing unit test to ensure error returns (via `false`
command) without timeouts are properly returned.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
This commit implements this corner case during osd-prepare job.
```
The encrypted block is not opened, this is an extreme corner case
The OSD deployment has been removed manually AND the node rebooted
So we need to re-open the block to re-hydrate the OSDInfo.
Handling this case would mean, writing the encryption key on a
temporary file, then call luksOpen to open the encrypted block and
then call ceph-volume to list against the opened encrypted block.
We don't implement this, yet and return an error.
```
When underlying PVC for osd are CSI provisioned, the encrypted device
is closed when PVC is unmounted due to osd pod being deleted.
Therefore, this may occur more frequently and needs to be handled.
This commit implements the fix for the same.
Signed-off-by: Rakshith R <rar@redhat.com>
Rook should update the RGW object store's period if the period doesn't
yet exist. This protects us from the case where the
'radosgw-admin period update --commit` command fails and the
CephObjectStore controller reconciles again.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
In ExtractExitCode, there is one error type that can be valid as a
pointer or not-as-a-pointer. Add a case to the type check for the
non-pointer condition.
Resolves#8280
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
We don't always get an error of the type "*exec.ExitError" so we must
validate the type before printing it otherwise the interface conversion
will fail.
Signed-off-by: Sébastien Han <seb@redhat.com>