Skip to content

[AI-7115] Add a request retry strategy to the async GitHub client - #24963

Draft
AAraKKe wants to merge 3 commits into
masterfrom
aarakke/AI-7115-github-client-retry
Draft

[AI-7115] Add a request retry strategy to the async GitHub client#24963
AAraKKe wants to merge 3 commits into
masterfrom
aarakke/AI-7115-github-client-retry

Conversation

@AAraKKe

@AAraKKe AAraKKe commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator

What does this PR do?

Adds a retry strategy to the async GitHub client for the failures that are not rate limiting, and separates it from the rate-limit handling that was already there.

Structure worth knowing before reading the diff:

  • Two layers, deliberately nested. _request (retries) wraps _rate_limited_request (today's loop, renamed). That order matters: each retry re-acquires the limiter, so it waits out any pause the governor is holding. Rate-limit responses stay owned by the inner layer and are never retried by the outer one.
  • retry.py describes, stamina executes. RetryPolicy is data: what to retry on, how many attempts, what backoff. No sleeping or backoff arithmetic of ours.
  • Defaults are chosen per endpoint by whether the request can be replayed, not by verb. Three mutating endpoints are idempotent and say so at their call site. Every method takes retry= to override, and policies compose.
  • A guard sits above any policy: auth failures, rate-limit responses, the limiter's give-up signal and redirects are never retried, however a caller configures things.
  • download_artifact retries as a pair. The signed URL expires, so the retry has to re-resolve the redirect rather than refetch a dead URL.
  • Config tunes the ladder only ([dispatcher.github_retries]). Widening what may be retried would make a duplicate side effect a setting.

ddev/src/ddev/utils/github_async/AGENTS.md documents the layer boundary so the next change lands in the right one.

Motivation

Closes AI-7115.

Dispatcher runs for hours and makes thousands of GitHub calls, and until now any failure that was not rate limiting failed on the first attempt. That gives a single blip more power than it should have. TaskTestRunner polls get_workflow_run for the whole life of a batch inside a try/finally with no except, so one transient 500 aborts the batch, closes its check run as cancelled and throws away the results of every test in it. Other calls swallow the failure and quietly degrade instead: a failed list_workflow_jobs returns an empty job list, so job correlation silently loses data.

Both get worse as we scale up: more batches and more polling mean more chances to hit the one blip that costs a whole batch of test results. Retrying is also a precondition for trusting the run report, since a report that is missing jobs because of a dropped connection is worse than one that is late.

No task behaviour changes here. Retries only make those paths less likely to fire, and a failure that outlives the ladder surfaces exactly as it does today.

Notes for review

  • "Retry" already means re-running failed test jobs in Dispatcher (ExecutionState.RETRYING, BatchProgress.retrying_jobs). This is unrelated and only concerns HTTP requests.
  • test_no_retry_on_transport_error became test_the_rate_limit_layer_does_not_retry_a_transport_error and now calls _rate_limited_request. The property still holds for that layer, but at client level a GET transport error is now retried on purpose.
  • The artifact policy adds 403 because that is how an expired signed URL presents from the storage host. GitHub's own 403 arrives as GitHubAuthenticationError, which the guard refuses, so a real denial still fails immediately. Tested both ways.
  • Open question, no action taken: stamina installs a process-wide hook that logs every scheduled retry to the stamina logger, so retries are visible even with no logger injected. Turning it off is global and would also silence the unrelated stamina.retry in ddev/e2e/agent/docker.py, so I left it and documented it. Say the word if you want the client to be the only voice.

Review checklist (to be filled by reviewers)

  • Feature or bugfix MUST have appropriate tests (unit, integration, e2e)
  • Add qa/required if this PR needs QA validation, or qa/skip-qa if it does not. Exactly one of the two is required.
  • If you need to backport this PR to another branch, you can add the backport/<branch-name> label to the PR and it will automatically open a backport PR once this one is merged

- New retry.py: RetryPolicy plus composable predicates, executed by stamina.
- Split the two layers: _request retries, _rate_limited_request handles rate limits.
- Per-endpoint defaults by replay safety, overridable per call with retry=.
- Never follow or retry an unexpected redirect; report it with the endpoint.
- Retry the artifact redirect and signed download as a pair.
- Expose the limits through [dispatcher.github_retries].
@AAraKKe AAraKKe added the qa/skip-qa Automatically skip this PR for the next QA label Aug 24, 2026
@dd-octo-sts dd-octo-sts Bot added the ddev label Aug 24, 2026
@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Aug 24, 2026

Copy link
Copy Markdown

evalya-impact-summary

evalya impact analysis
Impact analysis: 0 selected, 0 skipped (of 0 test tasks)
Publish tasks:   1 (always emitted)
Diff (13 files):
  ddev/changelog.d/24963.added
  ddev/src/ddev/cli/ci/tests/dispatcher_config.py
  ddev/src/ddev/utils/github_async/AGENTS.md
  ddev/src/ddev/utils/github_async/__init__.py
  ddev/src/ddev/utils/github_async/client.py
  ddev/src/ddev/utils/github_async/retry.py
  ddev/src/ddev/utils/github_errors.py
  ddev/tests/cli/ci/tests/test_dispatcher_config.py
  ddev/tests/utils/github_async/conftest.py
  ddev/tests/utils/github_async/helpers.py
  ddev/tests/utils/github_async/test_download_artifact.py
  ddev/tests/utils/github_async/test_rate_limiting.py
  ddev/tests/utils/github_async/test_retry.py

Debug a specific task: evalya plan impact --path <path> --task <task>

Learn more about CI impact filtering

@datadog-prod-us1-6

datadog-prod-us1-6 Bot commented Aug 24, 2026

Copy link
Copy Markdown

Tests  Code Coverage

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🚧 4 tests that failed were ignored due to quarantine View in Datadog

🎯 Code Coverage (details)
Patch Coverage: 100.00%
Overall Coverage: 88.77%

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: f3b5071 | Docs | View more details | Give us feedback!

@AAraKKe

AAraKKe commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c86bea94d0

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +292 to +293
assert NO_RETRY.should_retry is never
assert NO_RETRY.attempts == 1

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Remove the implementation-only NO_RETRY test

These assertions verify the constant's internal predicate identity and configured field rather than any observable request behavior, and they duplicate test_a_caller_can_turn_retrying_off_for_one_call, which already proves that NO_RETRY prevents a failed request from being replayed. An equivalent implementation could therefore break this test without changing the contract; keep the boundary-level test instead.

AGENTS.md reference: AGENTS.md:L181-L185

Useful? React with 👍 / 👎.

- Move the retry config into dispatcher_config, next to the other config models.
- Group module constants at the top of retry.py and trim the comments.
- RetryPolicy is a plain class with a typed replace instead of a dataclass.
- Move the client-specific guard and the retry cause into the client module.
- Redact the query string from the artifact URL before logging it.
@dd-octo-sts

dd-octo-sts Bot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

Validation Report

All 21 validations passed.

Show details
Validation Description Status
agent-reqs Verify check versions match the Agent requirements file
ci Validate CI configuration and code coverage settings
codeowners Validate every integration has a CODEOWNERS entry
config Validate default configuration files against spec.yaml
dep Verify dependency pins are consistent and Agent-compatible
http Validate integrations use the HTTP wrapper correctly
imports Validate check imports do not use deprecated modules
integration-style Validate check code style conventions
jmx-metrics Validate JMX metrics definition files and config
labeler Validate PR labeler config matches integration directories
legacy-signature Validate no integration uses the legacy Agent check signature
license-headers Validate Python files have proper license headers
licenses Validate third-party license attribution list
metadata Validate metadata.csv metric definitions
models Validate configuration data models match spec.yaml
openmetrics Validate OpenMetrics integrations disable the metric limit
package Validate Python package metadata and naming
qa-label Validate the pull request declares whether it needs QA for the next Agent release
readmes Validate README files have required sections
saved-views Validate saved view JSON file structure and fields
version Validate version consistency between package and changelog

View full run

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ddev qa/skip-qa Automatically skip this PR for the next QA

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant