The keystone auth suite installs cert-manager and trust-manager with
helm --wait and then immediately applies Issuer, Certificate, and
Bundle resources. helm --wait returns when the webhook deployments
report ready, but the webhook service endpoints may not be programmed
on the apiserver's node yet, so the apply is rejected with 'failed
calling webhook ... connect: connection refused' and the suite fails
before any test runs. This is currently the dominant failure mode of
the keystone suite, seen on master and on PRs that cannot have caused
it (dependabot github-actions bumps, mergify backports).
Add utils.Eventually, a timeout-based polling primitive intended as
the shared replacement for hand-rolled retry loops in test code:
- cond is func(ctx) error rather than func() bool, so each failed
attempt logs its reason via t.Logf and the final timeout error wraps
the last attempt's error.
- cond runs on the calling goroutine and receives a context carrying
the overall deadline, so cooperative operations stop at the deadline
instead of overrunning it. It is never run on another goroutine: a
hung attempt would keep executing concurrently with its retry, and
require's FailNow is unsafe off the test goroutine.
- utils.AttemptTimeout decorates a cond with a per-attempt deadline
for operations that can hang but would succeed if canceled and
retried. The bound is cooperative; conds that cannot honor a context
must be bounded at the operation level instead.
Expose it as a K8sHelper.ApplyWithRetry method that retries kubectl
apply until the webhooks accept connections, and use it for the
keystone setup applies instead of failing the suite on the first
attempt.
Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>