mirror of
https://github.com/wolfSSL/wolfssl.git
synced 2026-08-16 23:51:37 +02:00
A workflow file GitHub cannot load does not fail loudly. Its runs end within 0s with zero jobs, no logs, no annotations and no check runs, and the workflow re-registers under its bare path instead of its `name:` field. Among the few hundred checks on a PR that reads as unrelated flake, so the coverage just disappears: os-check.yml was in this state on master for ten days in July 2026 before anyone noticed, and no open PR reported a problem the whole time. Add two guards. Pre-merge, check-workflows.py measures every `run:` step against GitHub's 21000 character cap and fails the build past it, with a warning from 18000 so a growing step is noticed while there is still runway. Sizes come from the parsed YAML, which is what the Actions service evaluates, so block-scalar indentation needs no guessing. It runs from check-source-text.yml over every workflow and composite action rather than only PR-changed files: the cap applies per file, the whole sweep takes well under a second, and a file can be pushed over the line by a change elsewhere in the PR. Note that this cap is enforced by the service and not by the workflow schema, so neither a YAML validator nor actionlint reports it. Post-merge, workflow-health.yml runs check-workflow-health.py daily and looks for the symptom rather than any particular cause, so a workflow that stops loading for a reason nobody anticipated is still caught. Two signals: an active workflow whose registered name equals its path, and a completed run that failed with zero jobs (prefiltered on created_at == updated_at, so only a handful need a jobs lookup). Against the live repository the first signal flags os-check.yml and nothing else across 107 workflows, and reports clean on wolfTPM and wolfMQTT. It exits 1 on a finding and 2 when the check could not be carried out at all, because a missing token and a broken workflow call for different responses. Findings go into a single reused issue rather than another red check that would blend into the noise: the body is rewritten on each run, a comment is posted only when the set of affected workflows changes, and the issue closes itself once everything loads again. Finding that issue reliably turned out to be the fiddly part, and the approach here is the one that survived testing against a live repository. The issue is identified by both a dedicated label and its title, and looked up through the REST issues endpoint. Both halves of that identity matter: the label alone is a normal repository label that anyone can apply, and an adopted issue has its body overwritten and is then closed, so matching on the label alone would destroy a mislabelled issue. Searching by title instead is unusable, because search ignores --state and returns closed issues, which had the monitor re-closing an already closed issue on every clean run. `gh issue list` reads a GraphQL replica that can lag. The REST endpoint lags too, by about 2.4s for a newly created issue, so the lookup re-checks a few times before concluding nothing is open - without that, consecutive runs each open a duplicate, and a clean run right after an outage fails to close the issue it just opened. Verified against the commit that caused the outage: check-workflows.py fails onacff4d62a(21813 characters) and passes on its parentf5ace71dd, which it flags at 20662 - already inside the warning band, 338 characters short of breaking. The full issue lifecycle (open, repeat with no comment, comment on change, close, stay closed, reopen a fresh issue for a new outage) was exercised end to end against a live repository.