executeCommandWithTimeout wires a single bytes.Buffer to both cmd.Stdout
and cmd.Stderr and joins the command in a goroutine via cmd.Wait(). On
the timeout kill path it called cmd.Process.Kill() and then read that
buffer without waiting for cmd.Wait() to return. Kill() only signals the
process; it does not wait for os/exec's output-copier goroutines (joined
only by cmd.Wait) to finish, so reading the buffer there races with those
writers, which the race detector flags.
Only read the buffer once the copier goroutines have drained, signaled by
cmd.Wait() on the done channel. Because a killed process can leave an
orphaned descendant holding the output pipe open -- a D-state
cryptsetup/dmsetup child, or the downstream of a "sh -c '... | head'"
pipeline -- cmd.Wait() may never return, so bound the drain with a short
grace period and give up on the captured output rather than blocking the
caller forever. This keeps the bounded return the timeout path is meant
to guarantee. On a failed Kill() the process may still be writing, so
return without reading the buffer.
Add a regression test that drives the kill path with a child that ignores
SIGINT while writing to stdout; it fails under -race before this change.
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Fix duplicate words, incorrect articles (a/an), it's/its, and other small
grammar mistakes in Go comments and user-facing messages across pkg/, cmd/,
and tests/.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
Several godoc comments led with a stale or incorrect identifier, left
over from renames, exported/unexported changes, copy-paste between
sibling declarations, or plain typos. As a result the documented name no
longer matched the function, method, type, or var it describes. Correct
each leading word to the name of the declaration it documents.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
WriteFile and PathToProjectRoot had no callers anywhere in the repo
(verified with deadcode, staticcheck U1000, and repo-wide grep, tests
included). WriteFile has long been superseded by direct os.WriteFile
usage. Drop the now-unused bytes, fmt, and runtime imports; the file's
remaining helpers (WriteFileToLog, CreateTempFile) are still in use.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
kmod.go was entirely dead: IsBuiltinKernelModule, LoadKernelModule,
CheckKernelModuleParam, and the getKernelVersion helper had no callers
anywhere in the repo (verified with grep, deadcode, and staticcheck).
Remove the file.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
When writing to the operator log, where possible
the namespace of the resource on which the operator
is working will be logged as a standard prefix
to more easily troubleshoot what resource is being reconciled
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Creating the rgw admin user is taking on the order of two
minutes instead of the expected seconds, thus the command
was timing out and cause the object store user creation
to fail. The increased timeout allows the admin user creation
to complete, and other users to be created as well.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Refactored multiple conditional if-else blocks into switch statements to improve readability and maintainability. Also removed outdated staticcheck rule comments (QF1002, QF1003) from .golangci.yaml.
Signed-off-by: Carlos Barria <cbarria@yahoo.com>
When a rados namespace is deleted, it cannot be purged
until its images and snapshots are deleted. Now a condition
is being added to the rados namespace CR status so the user
does not need to read the operator log to discover the cause.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
When the blockpool is deleted and there are images or
snapshots in the pool, the deletion is blocked by the
finalizer until they are all removed. The condition added
to the cephblockpool status only indicated if there were
radosnamespace dependencies. A new condition is now added
to describe if there are images or snapshots.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Update the unit test for ExecuteCommandWithTimeout(). This unit failed a
CI run. This seems to have been a race condition where the previous
`cat` command was returning quickly enough that it was available to the
`select` statment at the same time as the timeout.
Resolve this issue by replacing the `cat` command with `sleep`, which
will be guaranteed to take longer to return than the timeout, ensuring
this race doesn't happen.
While working on the test, I also notice that the 30 second timeouts for
other tests would be too long in the event the commands hang for any
reason, so lower these to help ensure the CI won't time out due to
unforeseen events.
Also add a missing unit test to ensure error returns (via `false`
command) without timeouts are properly returned.
Signed-off-by: Blaine Gardner <blaine.gardner@ibm.com>
The ci was using a pretty old version og golangci-lint.
This updates to the latest version.
Additionally, it silences some
gosec integer conversion overflow false positves
and fixes some real errors of this category
and string format errors found by golangci-lint, while at it.
Co-authored-by: Blaine Gardner <b.blaine.gardner@gmail.com>
Co-authored-by: Travis Nielsen <tnielsen@redhat.com>
Signed-off-by: Michael Adam <obnox@samba.org>
currently device class and device type
labels were clubbed with eachother
create a seperate label, as these two labels
can be seperate
Signed-off-by: parth-gr <partharora1010@gmail.com>
Due to multi-site Ceph object resources (realms, zonegroups, zones)
reconciliation CLI commands being proxied through CephCluster's Mgr pods
in Multus-based scenarios, return codes were being additionally wrapped
by K8S' command streams and weren't adequately handled in-code, causing
reconciliation of CephObject(Realm|ZoneGroup|Zone) resources to fail, as
an exit code of 0, instead of the expected syscall.ENOENT (2) was returned.
Fixes#12833
Signed-off-by: zer0def <zv0uzqbqnncivw0afjsx79jmd19shpps3r5f2vjc6yv@definedaszero.xyz>
Rook uses global or node-local device class configuration if device-level
configurations doesn't exist. However, after introducing device-class
level resource configuration, rook set the default value of device class.
So global or node-level device class configuration has never been used
after that.
Closes: https://github.com/rook/rook/issues/11871
Closes: https://github.com/rook/rook/issues/11826
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
few functions got change as they were deprecated
for ex: ioutil.Readfile change to os.Readfile
ioutil.TempFile change to os.CreateTemp
And fixed golang-ci-lint-issues
Signed-off-by: parth-gr <paarora@redhat.com>
This commit implements this corner case during osd-prepare job.
```
The encrypted block is not opened, this is an extreme corner case
The OSD deployment has been removed manually AND the node rebooted
So we need to re-open the block to re-hydrate the OSDInfo.
Handling this case would mean, writing the encryption key on a
temporary file, then call luksOpen to open the encrypted block and
then call ceph-volume to list against the opened encrypted block.
We don't implement this, yet and return an error.
```
When underlying PVC for osd are CSI provisioned, the encrypted device
is closed when PVC is unmounted due to osd pod being deleted.
Therefore, this may occur more frequently and needs to be handled.
This commit implements the fix for the same.
Signed-off-by: Rakshith R <rar@redhat.com>
A new variable is added to rook-ceph-operator-config
ConfigMap to allow using loop devices for osd.
This feature is intended to be used for testing purposes only.
Signed-off-by: Shinya Hayashi <shinya-hayashi@cybozu.co.jp>
The tool gofmt in go 1.19 requires certain formatting
in the comments section for better rendering.
Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
Rook supports raw mode OSD in host-based cluster. So we can also
support OSD on logical volume in this kind of cluster.
Logical volumes aren't picked by filters (i.e. `useAllDevices: true`
and `device{Path,}Filter` to avoid unwanted LV consumption on upgrade.
Closes: https://github.com/rook/rook/issues/2047
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
- disk UUID should only be got for "disk" type device.
- It's not necessary to fail prepare job when failing to get disk UUID
- We should consider that `sgdisk` reports UUID even if there is no GPT.
Closes: https://github.com/rook/rook/issues/9948
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
In volume.go, remove if condition in `callCephVolume`
method where we always pass false.
In exec.go, removing `ExecuteCommandWithFile*` methods
which was not used anywhere.
Signed-off-by: subhamkrai <srai@redhat.com>
This is handling a tricky scenario where the OSD deployment is manually
removed and the OSD never reconvers. This is unlikely to happen, but
still OSD should be able to run after that action. Essentially after a
manual deletion, we need to run the prepare job again to re-hydrate the
OSD information so that the OSD deployment can be deployed.
On encryption, it is a little bit tricky since ceph-volume list again
the main block won't return anything, so we need to target the encrypted
block to list.
There is another case this PR does not handle, which is the removal of
the OSD deployment and then the node is restarted. This means that the
encrypted container is not opened anymore. However, opening it requires
more work like writing the key on the filesystem (if not coming from the
Kubernete secret, eg,. KMS vault) and then run luksOpen. This is an
extreme corner case probably not worth worrying about for now.
Signed-off-by: Sébastien Han <seb@redhat.com>
rook command doesn't interpret `logtostderr` option. It's OK to just
remove this option because `capnslog` outputs all logs to stdout
by default. It's better to keep `AddGoFlagSet()` call because
some libraries might define their own flags with Go's `flag` package.
Signed-off-by: Satoru Takeuchi <satoru.takeuchi@gmail.com>
The previous exit 32 check for loop device is 5 years old. Also, if the
device cannot be read it will be skipped anyway so let's report the
error and not hide it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Create a new log level for Rook that is hidden from users. This is the
most verbose log level, and it is the level developers would like to use
to get debug logs that are important for debugging but that could leak
senstivie information like credentials in production use.
If a user sets their debug level to "TRACE", they will merely get
"DEBUG" level logs. Only if they set "TRACE_INSECURE" will they get
trace logs, and those are likely to include insecure information. Rook
tries very hard not to leak sensitive information in logs even with
verbose "DEBUG" logs.
Resolves#8778
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
Rook should update the RGW object store's period if the period doesn't
yet exist. This protects us from the case where the
'radosgw-admin period update --commit` command fails and the
CephObjectStore controller reconciles again.
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
In ExtractExitCode, there is one error type that can be valid as a
pointer or not-as-a-pointer. Add a case to the type check for the
non-pointer condition.
Resolves#8280
Signed-off-by: Blaine Gardner <blaine.gardner@redhat.com>
This commit is a large refactor on how the operator starts, stops and
how it starts various sub-components such as the ceph-csi driver. It
also refines the way we cancel orchestrations. We don't use breakpoints
anymore but send our self a SIGUP to reload our controller runtime
manager.
The reload will happen under different circonstances like:
* a new adminission controller secret is created/deleted/changed
* a CephCluster CR is edited
As mentioned earlier, the csi driver now has its own controller, just
like flex. It reacts to change in the operator config map for particular
ROOK_CSI_ fields.
A second new controller for the operator's general config has been
created, it manages:
* the logging level
* the ceph CLI command timeout
* the discovery daemon
The operator reacts much more rapidly to cancellation events by stopping
the manager's context and reloading it.
Signed-off-by: Sébastien Han <seb@redhat.com>
Sometimes the error does not tell much, so as `exit status 1` and
printing the output along returning the error is useful.
For instance, I saw a job failing with no osd and the prepare job had
those lines:
```
exec: Running command: lsblk /dev/sdb1 --bytes --nodeps --pairs ....
inventory: skipping device "sdb1". exit status 1
```
We need to understand more about the lsblk issue.
Signed-off-by: Sébastien Han <seb@redhat.com>