Files
my-rook-config/pkg/operator/ceph/cluster
Anas Khan f79696157d mon: fix data race on failedMonSchedule during scheduling
assignMons schedules each mon in its own goroutine and shares a
failedMonSchedule bool to record whether any of them failed. The
resultLock mutex already guards the shared c.mapping.Schedule map
write, but the three failedMonSchedule = true assignments were done
without holding the lock. When two or more mons fail to schedule at
the same time (for example when waitForMonitorScheduling errors, the
node choice is nil, or getNodeInfoFromNode fails) their goroutines
write the flag concurrently, which is a write-write data race.

Guard the failedMonSchedule writes with the existing resultLock, the
same lock the goroutines already use for the map update. The post-Wait
read is left as is since the WaitGroup orders it after every write.

Add a regression test that fails scheduling for multiple mons at once
and confirms assignMons returns an error; run it with -race to catch
the race.

Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
(cherry picked from commit 0bbd5de84c)
2026-07-21 16:25:00 +00:00
..
2026-05-20 15:03:03 +00:00
2025-08-27 14:47:45 -06:00
2026-05-20 15:03:03 +00:00
2025-03-26 10:41:48 -07:00