Skip to content

oci: Retry transient registry failures when pulling - #407

Draft
cgwalters-bot wants to merge 4 commits into
composefs:mainfrom
cgwalters-forge:bot/oci-pull-retries
Draft

cgwalters-bot wants to merge 4 commits into
composefs:mainfrom
cgwalters-forge:bot/oci-pull-retries

Conversation

@cgwalters-bot

Copy link
Copy Markdown
Contributor

Pulls through skopeo have no retries at all, so a single 503 or dropped connection from e.g. quay.io fails the whole pull. That's a regular source of CI flakes for bootc (bootc-dev/bootc#2177). The proxy doesn't retry for us: in podman the retry loop lives in the caller (c/common pkg/retry), so this adds one here.

The proxy decides which errors are worth retrying. Since skopeo 1.19 it tags each failed reply with c/common's IsErrorRetryable() verdict, the same check podman's retry loop uses, and cgwalters-forge/containers-image-proxy-rs#2 exposes that as Error::is_retryable(). composefs-rs does no error-text matching of its own. With an older skopeo nothing gets retried, same as today. Only proxy errors count, so a local failure to import verified data is never retried.

Retries happen per proxy operation: opening the image, fetching the manifest, the config, and each layer or delta blob. That way a failed layer doesn't refetch the others. The schedule is podman's default, 3 retries after 1s, 2s and 4s, and a fixed delay can be set like --retry-delay. Only docker:// pulls retry. PullOptions gains a public retry: RetryPolicy field to tune or disable it. Since PullOptions isn't #[non_exhaustive], code that builds it with a struct literal and no ..Default::default() needs updating. cfsctl oci pull --retry N exposes the count (--retry 0 fails fast); the varlink API keeps the default. Note that pull_image() now retries by default too. Each attempt restarts the blob transfer, so a layer's Started progress event is sent again per attempt, and cfsctl replaces that layer's bar.

The first commit makes cfsctl print progress messages when stderr isn't a terminal. indicatif drops them otherwise, so retry warnings would never show up in CI logs. It could land on its own.

Before merging: the second commit points containers-image-proxy at the forge branch through [patch.crates-io]. The plan is to land cgwalters-forge/containers-image-proxy-rs#2 and release it as 0.11.1 first, then replace the [patch] with a version bump. The retry loop stays in composefs-rs for now; bootc already has its own, and switching bootc's classifier to is_retryable() is a separate follow-up (bootc-dev/bootc#2466).

The new integration test serves an OCI layout from a small in-process registry. The registry answers the first manifest request with a 503 and hangs up on the first request for each blob. The test checks that cfsctl oci pull docker://... recovers and yields the pinned image ID, and that --retry 0 fails. It goes through the real skopeo, so it also checks that skopeo classifies both failures as retryable. That classification needs skopeo 1.19, so the test parses skopeo --version and skips with older versions (upstream's ubuntu-24.04 runners have 1.13.3). The skopeo check is now one shared helper in the integration tests' main.rs; before, old_format.rs had its own copy.

Testing ran on a 16-core devspace at head 39b7496, with proxy-rs pulled from the forge branch. Two containers were used:

  • quay.io/fedora/fedora:latest (rustc 1.98, skopeo 1.22.3): just fmt-check, just clippy and just check-feature-combos were clean, and cargo test -p composefs-oci -p composefs-ctl passed (174 composefs-oci tests, including the retry and fetch_layer unit tests). just test-integration passed 112, including test_pull_retries_transient_registry_errors and the new test_parse_skopeo_version. It failed test_ostree_pull_local_all_modes and test_varlink_open_repository_invalid_spec, and main fails the same two in that container.
  • ubuntu:24.04 (skopeo 1.13.3, rust stable), as in upstream's smoke job: the retry test reports skopeo (1, 13, 3) does not classify retryable errors (needs (1, 19, 0)), skipping registry retry test and passes, and test_parse_skopeo_version passes. I stopped the full unprivileged suite there after it sat for 30 minutes with idle cfsctl processes. I didn't dig into that, and main wasn't run there for comparison.

The fs-verity-dependent unit tests, just check-fuzz and the privileged VM tests weren't run for this revision. The earlier revision passed check-fuzz.

Closes #348

The Signed-off-by: Colin Walters <walters@verbum.org> on these commits was added on cgwalters's approval of the review draft: cgwalters-forge#1 (review)

Generated-by: https://github.com/cgwalters/#llms

indicatif hides progress bars when stderr is not a terminal, and
MultiProgress::println() then silently discards the message. So status
messages from a pull never show up in CI logs, which is where the
upcoming retry warnings for flaky registries matter most.

Note this means scripted (non-tty) callers now also see messages like
"Fetching config ..." on stderr.

Generated-by: AI
Signed-off-by: Colin Walters <walters@verbum.org>
Prep for retrying pulls, which needs the proxy's classification of
retryable errors. This is only until a containers-image-proxy release
includes it; then it becomes a plain version bump.

Generated-by: AI
Signed-off-by: Colin Walters <walters@verbum.org>
Pulls through skopeo have no retries at all, so a single 503 or dropped
connection from e.g. quay.io fails the whole pull; this is a regular
source of CI flakes for bootc. The proxy doesn't retry for us: in
podman the retry loop lives in the caller (c/common pkg/retry), so do
the same here.

Which errors are worth retrying is the proxy's call: since skopeo 1.19
it classifies each failure with c/common's IsErrorRetryable(), the
heuristic podman's retry loop uses, and containers-image-proxy exposes
that as Error::is_retryable(). With an older skopeo nothing is retried.
Only errors from the proxy are considered; a local failure to import
verified data is never retried.

Retries happen per proxy operation (opening the image, fetching the
manifest, the config, and each layer or delta blob), so a failure in
one layer doesn't refetch the others. The schedule is podman's default:
3 retries after 1s, 2s and 4s, or a fixed delay like its --retry-delay;
PullOptions::retry tunes or disables it. Only registry (docker://)
pulls retry. Note pull_image() now retries with the default policy too.

A layer is still only registered once the proxy has verified the whole
blob, and each attempt starts a fresh request. We now also wait for
the proxy's verdict when the import itself failed, since a broken
transfer is then usually the root cause and decides whether to retry;
the import error is attached to it as context. Since each attempt
restarts the transfer, a layer's Started event is sent again per
attempt, and cfsctl replaces that layer's progress bar.

Retries are reported as progress messages, which is how cfsctl shows
them, and only logged at debug level so they don't appear twice.

Generated-by: AI
Closes: composefs#348
Signed-off-by: Colin Walters <walters@verbum.org>
Expose the retry count for registry pulls, so it can be raised for
flaky environments or set to 0 to fail fast. The varlink API keeps the
default policy for now.

The integration test serves a local OCI layout from a minimal
in-process registry that answers the first manifest request with a 503
and hangs up on the first request for each blob (skopeo retries
neither itself), and checks that the pull recovers and yields the
pinned image ID, while --retry 0 fails. This goes through the real
skopeo proxy, so it also checks that skopeo classifies both failures
as retryable. That classification needs skopeo 1.19, so the test is
skipped with older ones (e.g. Ubuntu 24.04 has 1.13); the skopeo
version check is now shared by all integration tests.

Generated-by: AI
Signed-off-by: Colin Walters <walters@verbum.org>
@cgwalters-bot

Copy link
Copy Markdown
Contributor Author

Rebased onto main; 4 commits, no content change.

Generated-by: https://github.com/cgwalters/#llms

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

oci: Implement fetching retries

1 participant