oci: Retry transient registry failures when pulling - #407
Draft
cgwalters-bot wants to merge 4 commits into
Draft
cgwalters-bot wants to merge 4 commits into
cgwalters-bot wants to merge 4 commits into
Conversation
This was referenced Sep 29, 2026
indicatif hides progress bars when stderr is not a terminal, and MultiProgress::println() then silently discards the message. So status messages from a pull never show up in CI logs, which is where the upcoming retry warnings for flaky registries matter most. Note this means scripted (non-tty) callers now also see messages like "Fetching config ..." on stderr. Generated-by: AI Signed-off-by: Colin Walters <walters@verbum.org>
Prep for retrying pulls, which needs the proxy's classification of retryable errors. This is only until a containers-image-proxy release includes it; then it becomes a plain version bump. Generated-by: AI Signed-off-by: Colin Walters <walters@verbum.org>
Pulls through skopeo have no retries at all, so a single 503 or dropped connection from e.g. quay.io fails the whole pull; this is a regular source of CI flakes for bootc. The proxy doesn't retry for us: in podman the retry loop lives in the caller (c/common pkg/retry), so do the same here. Which errors are worth retrying is the proxy's call: since skopeo 1.19 it classifies each failure with c/common's IsErrorRetryable(), the heuristic podman's retry loop uses, and containers-image-proxy exposes that as Error::is_retryable(). With an older skopeo nothing is retried. Only errors from the proxy are considered; a local failure to import verified data is never retried. Retries happen per proxy operation (opening the image, fetching the manifest, the config, and each layer or delta blob), so a failure in one layer doesn't refetch the others. The schedule is podman's default: 3 retries after 1s, 2s and 4s, or a fixed delay like its --retry-delay; PullOptions::retry tunes or disables it. Only registry (docker://) pulls retry. Note pull_image() now retries with the default policy too. A layer is still only registered once the proxy has verified the whole blob, and each attempt starts a fresh request. We now also wait for the proxy's verdict when the import itself failed, since a broken transfer is then usually the root cause and decides whether to retry; the import error is attached to it as context. Since each attempt restarts the transfer, a layer's Started event is sent again per attempt, and cfsctl replaces that layer's progress bar. Retries are reported as progress messages, which is how cfsctl shows them, and only logged at debug level so they don't appear twice. Generated-by: AI Closes: composefs#348 Signed-off-by: Colin Walters <walters@verbum.org>
Expose the retry count for registry pulls, so it can be raised for flaky environments or set to 0 to fail fast. The varlink API keeps the default policy for now. The integration test serves a local OCI layout from a minimal in-process registry that answers the first manifest request with a 503 and hangs up on the first request for each blob (skopeo retries neither itself), and checks that the pull recovers and yields the pinned image ID, while --retry 0 fails. This goes through the real skopeo proxy, so it also checks that skopeo classifies both failures as retryable. That classification needs skopeo 1.19, so the test is skipped with older ones (e.g. Ubuntu 24.04 has 1.13); the skopeo version check is now shared by all integration tests. Generated-by: AI Signed-off-by: Colin Walters <walters@verbum.org>
Contributor
Author
|
Rebased onto main; 4 commits, no content change. Generated-by: https://github.com/cgwalters/#llms |
cgwalters-bot
force-pushed
the
bot/oci-pull-retries
branch
from
October 2, 2026 12:14
a594bfe to
1257fa2
Compare
This was referenced Oct 2, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pulls through skopeo have no retries at all, so a single 503 or dropped connection from e.g. quay.io fails the whole pull. That's a regular source of CI flakes for bootc (bootc-dev/bootc#2177). The proxy doesn't retry for us: in podman the retry loop lives in the caller (c/common
pkg/retry), so this adds one here.The proxy decides which errors are worth retrying. Since skopeo 1.19 it tags each failed reply with c/common's
IsErrorRetryable()verdict, the same check podman's retry loop uses, and cgwalters-forge/containers-image-proxy-rs#2 exposes that asError::is_retryable(). composefs-rs does no error-text matching of its own. With an older skopeo nothing gets retried, same as today. Only proxy errors count, so a local failure to import verified data is never retried.Retries happen per proxy operation: opening the image, fetching the manifest, the config, and each layer or delta blob. That way a failed layer doesn't refetch the others. The schedule is podman's default, 3 retries after 1s, 2s and 4s, and a fixed delay can be set like
--retry-delay. Onlydocker://pulls retry.PullOptionsgains a publicretry: RetryPolicyfield to tune or disable it. SincePullOptionsisn't#[non_exhaustive], code that builds it with a struct literal and no..Default::default()needs updating.cfsctl oci pull --retry Nexposes the count (--retry 0fails fast); the varlink API keeps the default. Note thatpull_image()now retries by default too. Each attempt restarts the blob transfer, so a layer'sStartedprogress event is sent again per attempt, and cfsctl replaces that layer's bar.The first commit makes cfsctl print progress messages when stderr isn't a terminal. indicatif drops them otherwise, so retry warnings would never show up in CI logs. It could land on its own.
Before merging: the second commit points
containers-image-proxyat the forge branch through[patch.crates-io]. The plan is to land cgwalters-forge/containers-image-proxy-rs#2 and release it as 0.11.1 first, then replace the[patch]with a version bump. The retry loop stays in composefs-rs for now; bootc already has its own, and switching bootc's classifier tois_retryable()is a separate follow-up (bootc-dev/bootc#2466).The new integration test serves an OCI layout from a small in-process registry. The registry answers the first manifest request with a 503 and hangs up on the first request for each blob. The test checks that
cfsctl oci pull docker://...recovers and yields the pinned image ID, and that--retry 0fails. It goes through the real skopeo, so it also checks that skopeo classifies both failures as retryable. That classification needs skopeo 1.19, so the test parsesskopeo --versionand skips with older versions (upstream's ubuntu-24.04 runners have 1.13.3). The skopeo check is now one shared helper in the integration tests'main.rs; before,old_format.rshad its own copy.Testing ran on a 16-core devspace at head 39b7496, with proxy-rs pulled from the forge branch. Two containers were used:
quay.io/fedora/fedora:latest(rustc 1.98, skopeo 1.22.3):just fmt-check,just clippyandjust check-feature-comboswere clean, andcargo test -p composefs-oci -p composefs-ctlpassed (174 composefs-oci tests, including the retry andfetch_layerunit tests).just test-integrationpassed 112, includingtest_pull_retries_transient_registry_errorsand the newtest_parse_skopeo_version. It failedtest_ostree_pull_local_all_modesandtest_varlink_open_repository_invalid_spec, and main fails the same two in that container.ubuntu:24.04(skopeo 1.13.3, rust stable), as in upstream's smoke job: the retry test reportsskopeo (1, 13, 3) does not classify retryable errors (needs (1, 19, 0)), skipping registry retry testand passes, andtest_parse_skopeo_versionpasses. I stopped the full unprivileged suite there after it sat for 30 minutes with idlecfsctlprocesses. I didn't dig into that, and main wasn't run there for comparison.The fs-verity-dependent unit tests,
just check-fuzzand the privileged VM tests weren't run for this revision. The earlier revision passedcheck-fuzz.Closes #348
The
Signed-off-by: Colin Walters <walters@verbum.org>on these commits was added on cgwalters's approval of the review draft: cgwalters-forge#1 (review)Generated-by: https://github.com/cgwalters/#llms