Commit Graph
5329 Commits
Author SHA1 Message Date
Blaine Gardner 472c4f3f92 Merge pull request #18010 from jhoblitt/ss-17970
object: return an error when the multisite zone is not found
2026-07-30 13:50:40 -06:00
Blaine Gardner 908816c09b Merge pull request #18038 from parth-gr/offboard-accounts
rgw: force delete the rgw accounts by looking at the force annotation
2026-07-29 13:55:04 -06:00
Travis Nielsen d1a92df02b Merge pull request #17953 from cobaltcore-dev/osd-replacement-2-6
Osd replacement impl
2026-07-29 12:03:23 -06:00
Travis Nielsen ac0c3b45bc Merge pull request #18057 from taraasrita10/add-watcher-svg
external: add watcher to the svg controller for watching csi and clie…
2026-07-29 11:05:08 -06:00
Blaine Gardner db49154b8d Merge pull request #18012 from jhoblitt/ss-17964
core: fix data race reading exec output on timeout
2026-07-29 09:31:19 -06:00
Artem Torubarov c9b1bfa805 osd: use a separate annotation to fence osd replacement instead of existing label
osd-replacement was relying on exising ceph.rook.io/do-not-reconcile
label for fencing osd destroy process owned by osd health goroutine from
controller. However, this label can be used but other components like
rook krew maintenance plugin and cannot be owned by osd-replacement
process. Added a separate osd.rook.io/replace-in-progress annotation for
that purpose.

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-29 11:03:29 +02:00
parth-gr d6a0e9fd87 rgw: force delete the rgw accounts by looking at the force annotation
if the account cr has a force delete annotation, forcefully
remove the account, even if it contatins data in it, using the
purge data flag

Signed-off-by: parth-gr <partharora1010@gmail.com>
2026-07-29 13:03:06 +05:30
Travis Nielsen 6ed8f3efcd Merge pull request #18028 from taraasrita10/support_EC_with_MSR
pool: support CRUSH MSR rules for EC pools
2026-07-28 13:45:02 -06:00
Prabhala Tara Aasrita 715f38de91 pool: support CRUSH MSR rules for EC pools
Added failureDomains and osdsPerDomain fields to the
ErasureCodedSpec, which maps to the crush-num-failure-domains and
crush-osds-per-failure-domain Ceph EC profile parameters. This enables
creating EC pools that distribute chunks across fewer, larger hosts
without needing one host per data chunk.

Signed-off-by: Prabhala Tara Aasrita <taraprabhala@Prabhalas-MacBook-Pro.local>
2026-07-28 20:23:51 +05:30
Blaine Gardner d01d770a04 Merge pull request #18029 from raaizik/nfs-custom-port
nfs: support custom NFS server port
2026-07-28 08:49:20 -06:00
Prabhala Tara Aasrita a0c59bd7e8 external: add watcher to the svg controller for watching csi and client profile
Added a watcher to the SubVolumeGroup controller for watching csi and
client profile.

CLoses #18030

Signed-off-by: Prabhala Tara Aasrita <taraprabhala@Prabhalas-MacBook-Pro.local>
2026-07-28 11:55:11 +05:30
dependabot[bot] 6b126236b2 build(deps): bump the k8s-dependencies group across 1 directory with 5 updates
Bumps the k8s-dependencies group with 3 updates in the / directory: [k8s.io/api](https://github.com/kubernetes/api), [k8s.io/cli-runtime](https://github.com/kubernetes/cli-runtime) and [k8s.io/cloud-provider](https://github.com/kubernetes/cloud-provider).


Updates `k8s.io/api` from 0.36.2 to 0.36.3
- [Commits](https://github.com/kubernetes/api/compare/v0.36.2...v0.36.3)

Updates `k8s.io/apimachinery` from 0.36.2 to 0.36.3
- [Commits](https://github.com/kubernetes/apimachinery/compare/v0.36.2...v0.36.3)

Updates `k8s.io/cli-runtime` from 0.36.2 to 0.36.3
- [Commits](https://github.com/kubernetes/cli-runtime/compare/v0.36.2...v0.36.3)

Updates `k8s.io/client-go` from 0.36.2 to 0.36.3
- [Changelog](https://github.com/kubernetes/client-go/blob/master/CHANGELOG.md)
- [Commits](https://github.com/kubernetes/client-go/compare/v0.36.2...v0.36.3)

Updates `k8s.io/cloud-provider` from 0.36.2 to 0.36.3
- [Commits](https://github.com/kubernetes/cloud-provider/compare/v0.36.2...v0.36.3)

---
updated-dependencies:
- dependency-name: k8s.io/api
  dependency-version: 0.36.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: k8s-dependencies
- dependency-name: k8s.io/apimachinery
  dependency-version: 0.36.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: k8s-dependencies
- dependency-name: k8s.io/cli-runtime
  dependency-version: 0.36.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: k8s-dependencies
- dependency-name: k8s.io/client-go
  dependency-version: 0.36.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: k8s-dependencies
- dependency-name: k8s.io/cloud-provider
  dependency-version: 0.36.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: k8s-dependencies
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-27 21:12:56 +00:00
Travis Nielsen 72ec70d226 Merge pull request #18050 from rook/dependabot/go_modules/github-dependencies-2741939467
build(deps): bump the github-dependencies group with 6 updates
2026-07-27 15:10:57 -06:00
Artem Torubarov d18ff0da06 osd: detect in-use to empty device transition so a disk swap triggers a reconcile
The daemon only flagged a device change from in-use to empty, never the
reverse, so an in-use disk was never recorded as non-empty and a later
disk swap produced no ConfigMap update or reconcile. This is needed for
OSD replacement to work automatically: without it a destroyed OSD slot is
never reprovisioned after the disk is swapped. Compare the empty state in
both directions so the swap becomes a real transition that triggers the
reconcile.

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:44:01 +02:00
Artem Torubarov 0eb4735ee3 osd: device class matching for osd replacement
in case of multiple destroyed OSDs or multiple available
blank devices, Rook will try to match them by the same
device class if possible. If no matching DC, then DC of
destroyed OSD will be used

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Artem Torubarov 11ffd3b717 osd: use helper to check command status code
Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Artem Torubarov d7e777723f osd: implement osd replacement pdb exemption
Destroyed OSDs awaiting replacement from osd-replacement flow should
be ignored by managed PDB to avoid the following situation:
replaced OSD is already destroyed, its deployment is downscaled zero
and it awaits the physical swap (can take days). This situation does
not trigger blocking PDB in happy case because OSD does not holds PGs.
However, if PGs go unclean during that window for an unrelated reason,
the destroyed OSD gets counted as a down OSD and can grab the single
"draining failure domain" slot - putting blocking PDBs on every other FD.

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Artem Torubarov a0d215b428 osd: implement osd replacement health goroutine
A state machine for osd destroy phase of osd-replacement flow.
The state machine is implemented in OSD health goroutine, where each
step is handled on a separate tick. State machine is responsible
for draining, destroying OSD, and reserving its CRUSH position,
downscalling its deployment, and adding "readyForSwap" annotation.
Destroy phase does not zap the device on purpose. On disk signature
is kept to avoid automatic reprovisioning of destroyed OSD and to
detect physical swap.

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Artem Torubarov 514bd21552 osd: implement osd replacement controller
adds support of osd-replacement flow to cluster controller:
- predicate for osd deployment annotation
- validation of osd replacement annotation
- osd deployment recreation logic for replaced OSD

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Artem Torubarov c4727a0b21 osd: implement osd replacement crypto close job
defines a new entry point "close-encrypted-devices" in osd job.
The job is used by osd-replacement flow to close dm-crypt mappings
on host for destroyed encrypted OSDs. The job is owned by OSD
health goroutine implementing destroy phase of osd-replacement.

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
Artem Torubarov 887d533001 osd: implement osd replacement prepare-job
contains osd prepare job edits to support osd replacement flow:
- removes destroyed OSD IDs from node status CM to stop cluster
  controller from reprovisioning destroyed OSD
- adds provisioning logic for OSD replacement to match available
  blank devices to destroyed OSD slots

Signed-off-by: Artem Torubarov <artem.torubarov@sap.com>
2026-07-27 18:42:07 +02:00
dependabot[bot] 9a7005a870 build(deps): bump go.yaml.in/yaml/v3 from 3.0.4 to 3.0.5
Bumps [go.yaml.in/yaml/v3](https://github.com/yaml/go-yaml) from 3.0.4 to 3.0.5.
- [Commits](https://github.com/yaml/go-yaml/compare/v3.0.4...v3.0.5)

---
updated-dependencies:
- dependency-name: go.yaml.in/yaml/v3
  dependency-version: 3.0.5
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-27 12:50:37 +00:00
dependabot[bot] 03d5028ef9 build(deps): bump the github-dependencies group with 6 updates
Bumps the github-dependencies group with 6 updates:

| Package | From | To |
| --- | --- | --- |
| [github.com/aws/aws-sdk-go-v2](https://github.com/aws/aws-sdk-go-v2) | `1.42.1` | `1.43.0` |
| [github.com/aws/aws-sdk-go-v2/config](https://github.com/aws/aws-sdk-go-v2) | `1.32.30` | `1.32.31` |
| [github.com/aws/aws-sdk-go-v2/credentials](https://github.com/aws/aws-sdk-go-v2) | `1.19.29` | `1.19.30` |
| [github.com/aws/aws-sdk-go-v2/service/s3](https://github.com/aws/aws-sdk-go-v2) | `1.105.2` | `1.106.0` |
| [github.com/aws/aws-sdk-go-v2/service/sns](https://github.com/aws/aws-sdk-go-v2) | `1.41.1` | `1.42.0` |
| [github.com/go-logr/logr](https://github.com/go-logr/logr) | `1.4.3` | `1.4.4` |


Updates `github.com/aws/aws-sdk-go-v2` from 1.42.1 to 1.43.0
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/v1.42.1...v1.43.0)

Updates `github.com/aws/aws-sdk-go-v2/config` from 1.32.30 to 1.32.31
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/config/v1.32.30...config/v1.32.31)

Updates `github.com/aws/aws-sdk-go-v2/credentials` from 1.19.29 to 1.19.30
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/credentials/v1.19.29...credentials/v1.19.30)

Updates `github.com/aws/aws-sdk-go-v2/service/s3` from 1.105.2 to 1.106.0
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/service/s3/v1.105.2...service/s3/v1.106.0)

Updates `github.com/aws/aws-sdk-go-v2/service/sns` from 1.41.1 to 1.42.0
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/v1.41.1...v1.42.0)

Updates `github.com/go-logr/logr` from 1.4.3 to 1.4.4
- [Release notes](https://github.com/go-logr/logr/releases)
- [Changelog](https://github.com/go-logr/logr/blob/master/CHANGELOG.md)
- [Commits](https://github.com/go-logr/logr/compare/v1.4.3...v1.4.4)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2
  dependency-version: 1.43.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-dependencies
- dependency-name: github.com/aws/aws-sdk-go-v2/config
  dependency-version: 1.32.31
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-dependencies
- dependency-name: github.com/aws/aws-sdk-go-v2/credentials
  dependency-version: 1.19.30
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-dependencies
- dependency-name: github.com/aws/aws-sdk-go-v2/service/s3
  dependency-version: 1.106.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-dependencies
- dependency-name: github.com/aws/aws-sdk-go-v2/service/sns
  dependency-version: 1.42.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: github-dependencies
- dependency-name: github.com/go-logr/logr
  dependency-version: 1.4.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: github-dependencies
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-27 12:50:27 +00:00
raaizik b97296c47c nfs: add unit tests for custom NFS port
Add coverage for GetPort defaults and custom values, Service and probe
port wiring, Ganesha NFS_Port config, and create-or-update when
spec.server.port changes on an existing Service.

Signed-off-by: raaizik <raaizik@yahoo.com>
2026-07-27 13:54:57 +03:00
raaizik 5e36fa4bc3 nfs: support custom NFS server port
When NFS runs with host networking, port 2049 may already be in use.
Add CephNFS spec.server.port (default 2049), wire it through Ganesha
config, the operator Service, and probes, and regenerate CRDs.

Signed-off-by: raaizik <raaizik@yahoo.com>
2026-07-27 13:54:57 +03:00
Oded Viner f095a530c8 nvmeof: set nvmeof-meta application tag on .nvmeof pool
the .nvmeof pool was incorrectly tagged with the rbd
application and initialized as an rbd pool. add .nvmeof
to the built-in pool switch so it gets the nvmeof-meta
application tag instead, and skip rbd pool init.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-07-24 18:11:37 +03:00
Joshua Hoblitt 6fec883e62 Merge pull request #18013 from jhoblitt/ss-17966
osd: skip upgrades for OSDs pending migration when the last migration is incomplete
2026-07-23 07:52:49 -07:00
Travis Nielsen 9a52f468fc Merge pull request #17991 from chimanjain/migrate-yaml
build(deps): Migrate YAML library from gopkg.in/yaml.v2 to go.yaml.in/yaml/v3
2026-07-22 09:07:04 -06:00
Travis Nielsen 955f7c5428 Merge pull request #17785 from beyondcloud-co/fix/17761-duplicate-mgr-ips-in-endpointslice
fix(ceph): deduplicate external MGR endpoints in EndpointSlice
2026-07-22 09:05:31 -06:00
Chiman Jain 2034be1f37 build: migrate gopkg.in/yaml to go.yaml.in/yaml
Signed-off-by: Chiman Jain <chimanjain15@gmail.com>
2026-07-22 20:05:39 +05:30
Parth Arora 1bb5bdfb55 Merge pull request #17971 from parth-gr/watch-rns-svg
external: add watcher to the rns controler for watching csi and clien…
2026-07-22 19:57:09 +05:30
parth-gr 7f602f12b0 external: add watcher to the rns controler for watching csi and client profile
if the csi secret is updated or a client profile is updated
we need a re-reconile of rns
So added the watcher for it

Signed-off-by: parth-gr <partharora1010@gmail.com>
2026-07-22 19:40:12 +05:30
Travis Nielsen 8c397f6381 core: update the go text module
The go text package v0.37 has a vulnerability, thus update
to the latest version of the package.

Signed-off-by: Travis Nielsen <tnielsen@redhat.com>
2026-07-21 15:18:00 -06:00
Anas Khan 0bf6f84ff5 object: return an error when the multisite zone is not found
retrieveMultisiteZone is meant to gate the object-store reconcile on the
backing Ceph zone existing: it runs "radosgw-admin zone get" and, when
that fails, returns a non-nil error so the caller requeues. That gate is
dead code. The ENOENT check declares an inner err from exec.ExtractExitCode
that shadows the outer command error, and ExtractExitCode returns a nil
error for the ordinary exit failures radosgw-admin produces. Both the
ENOENT branch and the else branch then wrap that shadowed nil, and
errors.Wrapf(nil, ...) is nil, so the function returns
(waitForRequeueIfObjectStoreNotReady, nil). The caller only propagates the
requeue when the error is non-nil, so the requeue is dropped and reconcile
runs on.

The result is that a failed "zone get" no longer backs off. Reconcile
proceeds to stand up the object store anyway -- the RGW service, the
admin-ops endpoint, the deployment, and the pools radosgw scaffolds as it
comes up -- for a multisite store whose backing zone does not exist. This
is not the "normal multisite bootstrap" transient the original wording
suggested. getMultisiteResourceNames runs immediately before this and
already requeues until the CephObjectZone CR reports Ready, and the zone
controller marks it Ready only after it has created the Ceph zone, so a
healthy bootstrap never reaches this gate with a missing zone. That
CR-Ready gate, not this one, is what actually blocks bootstrap.

Where the dead gate does bite is the cases the CR status cannot cover:

- the Ceph zone deleted or renamed out of band while the CR still reads Ready
- a zone controller that reports Ready without leaving a usable zone behind
- any non-ENOENT "zone get" failure, e.g. a permission or connectivity
  error, which the else branch swallows the same way

The check was correct when fc579f4520 introduced it in 2020 with
exec.ExitStatus, which returns (code, ok) and leaves err unshadowed.
bd58790c31 ("ceph: proxy ceph commands when multus is configured", 2021)
swapped it to exec.ExtractExitCode as an unrelated drive-by, inverting the
contract and killing the gate. The sibling realm, zonegroup, and zone
controllers were left untouched and still use exec.ExitStatus today.

Restore exec.ExitStatus so the outer error is no longer shadowed, both
branches wrap the real command error, and the caller requeues until the
zone exists -- matching the sibling controllers.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-21 12:40:25 -07:00
Travis Nielsen ce05eff9e5 Merge pull request #17947 from OdedViner/fix_nvmeof_ci
nvmeof: use .nvmeof pool for gateway and remove pool field from CRD
2026-07-21 12:12:09 -06:00
Oded Viner ec502c3f80 nvmeof: use .nvmeof pool for gateway and remove pool field from CRD
Changes:
- Use a dedicated .nvmeof pool (via CephBlockPool CR named
  builtin-nvmeof) for the NVMe-oF gateway internal state.
- Use a separate nvmeof pool for the StorageClass data.
- Remove the pool field from the CephNVMeOFGateway CRD.
  The gateway now always uses the .nvmeof pool, hardcoded
  in the operator.
- Add .nvmeof to the allowed CephBlockPool name overrides.
- Create a production example nvmeof.yaml (instances: 2,
  replicas: 3) and a CI-only nvmeof-test.yaml (instances: 1,
  replicas: 1).
- Update documentation and CI test script accordingly.

Signed-off-by: Oded Viner <oviner@redhat.com>
2026-07-21 20:19:44 +03:00
Joshua Hoblitt 9312bbb350 Merge pull request #17980 from anxkhn/patch-10
core: parse subvolumegroup bytes_quota reported as "infinite"
2026-07-21 09:26:28 -07:00
Joshua Hoblitt a9bd575002 Merge pull request #17968 from anxkhn/patch-21
mon: fix data race on failedMonSchedule during scheduling
2026-07-21 09:24:24 -07:00
Anas Khan 7630d21e85 core: fix data race reading exec output on timeout
executeCommandWithTimeout wires a single bytes.Buffer to both cmd.Stdout
and cmd.Stderr and joins the command in a goroutine via cmd.Wait(). On
the timeout kill path it called cmd.Process.Kill() and then read that
buffer without waiting for cmd.Wait() to return. Kill() only signals the
process; it does not wait for os/exec's output-copier goroutines (joined
only by cmd.Wait) to finish, so reading the buffer there races with those
writers, which the race detector flags.

Only read the buffer once the copier goroutines have drained, signaled by
cmd.Wait() on the done channel. Because a killed process can leave an
orphaned descendant holding the output pipe open -- a D-state
cryptsetup/dmsetup child, or the downstream of a "sh -c '... | head'"
pipeline -- cmd.Wait() may never return, so bound the drain with a short
grace period and give up on the captured output rather than blocking the
caller forever. This keeps the bounded return the timeout path is meant
to guarantee. On a failed Kill() the process may still be writing, so
return without reading the buffer.

Add a regression test that drives the kill path with a child that ignores
SIGINT while writing to stdout; it fails under -race before this change.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-21 09:01:12 -07:00
Joshua Hoblitt 0e21738451 Merge pull request #18011 from jhoblitt/ss-17967
osd: archive crash on OSD removal
2026-07-21 08:39:35 -07:00
subham rai f5052c522c Merge pull request #18001 from rook/dependabot/go_modules/github-dependencies-6d2e00e983
build(deps): bump the github-dependencies group across 1 directory with 10 updates
2026-07-21 15:24:02 +05:30
Anas Khan 439b8c84d9 osd: skip upgrades for OSDs pending migration when the last migration is incomplete
startOSDMigration guards against upgrading OSDs while an OSD migration is
still in progress by returning the set of OSDs pending migration, which the
caller then removes from the update queue.

When the last migrated OSD's deployment has not been recreated yet,
isLastOSDMigrationComplete returns (false, nil). startOSDMigration then ran
return nil, errors.Wrapf(err, ...) with err already nil, and
errors.Wrapf(nil, ...) returns nil, so it returned (nil, nil), the same
value it uses to signal that no migration was requested. The caller
therefore skipped removing the pending OSDs from the update queue and could
upgrade and restart them mid-migration, which the pending-migration guard
was meant to prevent.

Return the freshly computed migration config in this case so the caller
still prunes the pending OSDs, without starting a new migration and without
aborting the reconcile. Aborting here would wedge the cluster because the
code that recreates the interrupted OSD runs later in Start, so an early
error return would leave that OSD deleted and block every later reconcile.
Letting the reconcile continue recreates the interrupted OSD and lets the
migration resume on a subsequent reconcile.

Also construct a non-nil error for the sibling unhealthy-PGs guard, which
wrapped a nil err in the same way.

Add a regression test asserting that startOSDMigration returns the pending
migration set, and does not delete another OSD deployment, while a previous
migration is still in progress.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-20 17:51:30 -07:00
Anas Khan 1630ad1049 osd: archive crash on OSD removal
archiveCrash is meant to silence the RECENT_CRASH health warning for a
removed OSD by archiving its crash, but the early-return guard was
inverted. GetCrashList unmarshals `ceph crash ls` into a non-nil slice on
every success, so `if crash != nil` was always true. The function always
logged "no ceph crash to silence" and the archive loop was unreachable, so
OSD removal never archived the crash and the warning lingered. Guard on
`len(crash) == 0` so only a genuinely empty list short-circuits.

Reviving the loop surfaced two latent bugs it had always masked. On a
non-empty list with no entry for the removed OSD the crash id stayed empty
and the code ran `ceph crash archive ""`, which the mgr rejects with EINVAL
and logs as a spurious error during routine OSD removal, so skip the
archive when no crash matches. The loop also broke on the first match, but
`ceph crash ls` is oldest-first and includes already-archived entries, so
an OSD with several crashes had only its oldest one archived while
RECENT_CRASH persisted; archive every matching entry instead of just the
first.

Extend TestArchiveCrash to cover all four cases: a single matching crash, a
list with several matching crashes that are all archived, an empty list,
and a non-empty list with no matching entry where nothing is archived.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-20 17:45:02 -07:00
Anas Khan 7f2a68fa91 object: return an error when reconciled rgw user has no keys
In createOrUpdateCephUser, when the desired user carries no explicit
keys the reconciler falls back to the live RGW user's keys. If that live
user also has zero keys the code intends to fail, but it built the error
with errors.Wrapf(err, ...) at a point where err is already nil (the
prior SetUserQuota error was handled and returned just above).
errors.Wrapf(nil, ...) returns nil, so the failure was swallowed and
createOrUpdateCephUser returned success on a user the operator itself
flagged as broken.

The reconcile then continued to generateCephUserSecret, which indexes
userConfig.Keys[0] to populate the Kubernetes secret and panicked on the
empty key slice. The reconciler's deferred RecoverAndLogException caught
and logged that panic, so the reconcile was abandoned before it reached
the Ready status update; and because the recovered Reconcile returns a
zero Result with a nil error, the request was not requeued either. The
user was left neither marked Ready nor retried.

Construct the error with errors.Errorf so the intended failure is
surfaced instead of being swallowed. Add a regression test that returns
a keyless live user and asserts a non-nil "no keys set" error.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
2026-07-20 17:41:21 -07:00
Joshua Hoblitt 7047206697 Merge pull request #17959 from anxkhn/patch-14
test: assert preferred anti-affinity count in testPodSpecPlacement
2026-07-20 17:31:12 -07:00
Joshua Hoblitt da490653a5 Merge pull request #17958 from anxkhn/patch-13
test: assert retained capacity LastUpdated in TestCephStatus
2026-07-20 17:29:56 -07:00
Joshua Hoblitt 101a9d4302 Merge pull request #17957 from anxkhn/patch-8
csi: fix compareJSON to compare actual output
2026-07-20 17:28:29 -07:00
Joshua Hoblitt 1cde80c6bb Merge pull request #17965 from anxkhn/patch-20
csi: do not log peer bootstrap token secret
2026-07-20 17:24:25 -07:00
Joshua Hoblitt 0f3d6b0941 Merge pull request #17963 from anxkhn/patch-19
mgr: use dashboard port 1024 as requested
2026-07-20 16:23:06 -07:00
Joshua Hoblitt 787d5a5878 Merge pull request #17961 from anxkhn/patch-7
docs: fix mirror Placement comments naming rgw pods
2026-07-20 14:58:20 -07:00