forked from rook/rook
The keystone auth suite installs cert-manager and trust-manager with helm --wait and then immediately applies Issuer, Certificate, and Bundle resources. helm --wait returns when the webhook deployments report ready, but the webhook service endpoints may not be programmed on the apiserver's node yet, so the apply is rejected with 'failed calling webhook ... connect: connection refused' and the suite fails before any test runs. This is currently the dominant failure mode of the keystone suite, seen on master and on PRs that cannot have caused it (dependabot github-actions bumps, mergify backports). Add utils.Eventually, a timeout-based polling primitive intended as the shared replacement for hand-rolled retry loops in test code: - cond is func(ctx) error rather than func() bool, so each failed attempt logs its reason via t.Logf and the final timeout error wraps the last attempt's error. - cond runs on the calling goroutine and receives a context carrying the overall deadline, so cooperative operations stop at the deadline instead of overrunning it. It is never run on another goroutine: a hung attempt would keep executing concurrently with its retry, and require's FailNow is unsafe off the test goroutine. - utils.AttemptTimeout decorates a cond with a per-attempt deadline for operations that can hang but would succeed if canceled and retried. The bound is cooperative; conds that cannot honor a context must be bounded at the operation level instead. Expose it as a K8sHelper.ApplyWithRetry method that retries kubectl apply until the webhooks accept connections, and use it for the keystone setup applies instead of failing the suite on the first attempt. Signed-off-by: Joshua Hoblitt <josh@hoblitt.com>
94 lines
3.5 KiB
Go
94 lines
3.5 KiB
Go
/*
|
|
Copyright 2026 The Rook Authors. All rights reserved.
|
|
|
|
Licensed under the Apache License, Version 2.0 (the "License");
|
|
you may not use this file except in compliance with the License.
|
|
You may obtain a copy of the License at
|
|
|
|
http://www.apache.org/licenses/LICENSE-2.0
|
|
|
|
Unless required by applicable law or agreed to in writing, software
|
|
distributed under the License is distributed on an "AS IS" BASIS,
|
|
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
See the License for the specific language governing permissions and
|
|
limitations under the License.
|
|
*/
|
|
|
|
package utils
|
|
|
|
import (
|
|
"context"
|
|
"fmt"
|
|
"testing"
|
|
"time"
|
|
)
|
|
|
|
// eventuallyPollInterval is the cadence of Eventually. It is deliberately not
|
|
// a parameter: consumers poll cheap idempotent operations (kubectl applies,
|
|
// rgw admin / S3 reads), and uniform call sites beat another knob. Add a
|
|
// variant if a real need for a different cadence ever appears.
|
|
const eventuallyPollInterval = time.Second
|
|
|
|
// Eventually polls cond on the calling goroutine every second until it
|
|
// returns nil, the timeout elapses, or ctx is canceled. Each failed attempt
|
|
// is logged via t.Logf so CI output shows liveness attributed to the right
|
|
// (sub)test, and the last attempt's error is wrapped into the returned error.
|
|
//
|
|
// The context passed to cond is derived from ctx with the overall timeout
|
|
// applied, so a cond that honors it stops at the deadline instead of
|
|
// overrunning it. To additionally bound each attempt, wrap cond with
|
|
// AttemptTimeout.
|
|
//
|
|
// cond must not call assert or require: report transient failures by
|
|
// returning an error, and make assertions on captured state after the wait.
|
|
// Eventually never runs cond on another goroutine — a hung attempt would keep
|
|
// executing concurrently with its retry, and require's FailNow is unsafe off
|
|
// the test goroutine (which is also why testify's assert.Eventually is not
|
|
// used here).
|
|
func Eventually(ctx context.Context, t *testing.T, timeout time.Duration, desc string, cond func(context.Context) error) error {
|
|
t.Helper()
|
|
|
|
loopCtx, cancel := context.WithTimeout(ctx, timeout)
|
|
defer cancel()
|
|
|
|
for attempt := 1; ; attempt++ {
|
|
err := cond(loopCtx)
|
|
if err == nil {
|
|
return nil
|
|
}
|
|
t.Logf("waiting for %s (attempt %d): %v", desc, attempt, err)
|
|
|
|
select {
|
|
case <-loopCtx.Done():
|
|
if cerr := ctx.Err(); cerr != nil {
|
|
return fmt.Errorf("canceled while waiting for %s: %w (last error: %v)", desc, cerr, err)
|
|
}
|
|
return fmt.Errorf("condition not met within %s: %s: %w", timeout, desc, err)
|
|
case <-time.After(eventuallyPollInterval):
|
|
}
|
|
}
|
|
}
|
|
|
|
// AttemptTimeout bounds each attempt of cond with its own deadline, for
|
|
// operations that can hang but would succeed if canceled and retried. The
|
|
// bound is cooperative: cond only stops early if it honors the context it is
|
|
// given (requests built with http.NewRequestWithContext, exec.CommandContext,
|
|
// SDK *WithContext calls, ...). A cond that ignores its context is unaffected
|
|
// by AttemptTimeout:
|
|
//
|
|
// // WRONG: http.Get ignores ctx, so nothing cancels a hung fetch
|
|
// utils.AttemptTimeout(10*time.Second, func(ctx context.Context) error {
|
|
// _, err := http.Get(url)
|
|
// return err
|
|
// })
|
|
//
|
|
// Operations that cannot honor a context must be bounded at the operation
|
|
// level instead (e.g. http.Client{Timeout: ...}).
|
|
func AttemptTimeout(d time.Duration, cond func(context.Context) error) func(context.Context) error {
|
|
return func(ctx context.Context) error {
|
|
attemptCtx, cancel := context.WithTimeout(ctx, d)
|
|
defer cancel()
|
|
return cond(attemptCtx)
|
|
}
|
|
}
|