Skip to content

fix(api-client): retry rate-limited requests after Retry-After - #388

Merged
gregberge merged 1 commit into
mainfrom
claude/modest-margulis-dbf31e
Oct 2, 2026
Merged

gregberge merged 1 commit into
mainfrom
claude/modest-margulis-dbf31e

Conversation

@gregberge

Copy link
Copy Markdown
Member

Description

What

  • apiFetch retries a 429 Too Many Requests after waiting for its Retry-After, up to 5 times, waiting at most 60s each time whatever the header asks.
  • When Retry-After is missing or invalid (anything but a whole number of seconds, which is what the API sends), it backs off exponentially from minTimeout: 1s, 2s, 4s, …
  • 429 retries are counted apart from the 3 server-error retries: they don't use them up, and a server-error retry doesn't reset the 429 count. Each retry replays the body, keeps x-argos-request-id and sends the next x-argos-retry-attempt, counted across both kinds of retry.
  • Aborting the request's signal ends the wait at once, rejecting with the signal's reason. The discarded 429's body is cancelled before the wait so the connection isn't held.
  • Once the retries run out, the last 429 response is returned unchanged, so callers still report the same APIError. The 5xx path is unchanged.

Why

apiFetch returned any 429 as is, so one throttled request failed the whole upload: upload() threw and the Playwright/Vitest reporters printed ❌ Error while creating the Argos build — APIError: HTTP 429 Too Many Requests: Too many requests, please try again later. A customer hit this on sharded CI runs.

The API's rate limiter (express-rate-limit with draft-8 headers and a Redis store) sends Retry-After: <seconds> on every 429, and counts requests in fixed 5-minute windows that start at the first request. A burst from an otherwise idle IP therefore usually gets a Retry-After close to the full 300s. Five waits of at most a minute is the smallest bound that still sees a request through a whole window. In exchange, rate limiting can hold a request up for 5 minutes at most, so CI can't hang on it.

This goes with the backend moving the SDK's CI upload endpoints to their own, larger rate-limit budget (argos-ci/argos, apps/backend/src/web/api/v2.ts): a burst over a limit should delay an upload rather than fail it.

Type of changes

Bug fix (bug label).

Checklist

  • I have read the CONTRIBUTING doc (there is no CONTRIBUTING doc in this repo)
  • The commits message follows the Conventional Commits' policy
  • Lint and unit tests pass locally
  • I have added tests if needed

Optional checks:

  • My changes requires a change to the documentation
  • I have updated the documentation accordingly

Further comments

Testing

  • 10 new cases in packages/api-client/src/fetch.test.ts, all on fake timers: the retry once Retry-After elapses (POST body replayed), the 60s cap, the fallback for a missing, empty, non-numeric or negative header, the last 429 returned after 5 retries (6 requests, 5 minutes in all), counting apart from server-error retries (headers checked), the 429 count surviving a server-error retry, and an abort during the wait. They all fail on the previous implementation, and breaking each part (cap, retry count, abort, fallback, counters) fails the matching test.
  • Ran the built client against a local Express server using the backend's express-rate-limit setup: a request over the limit got Retry-After: 3, waited 3s and succeeded; a request that was always rejected gave up after 6 attempts with the same HTTP 429 Too Many Requests: Too many requests, please try again later. message.
  • pnpm run test (all packages), pnpm run lint, pnpm run check-types and pnpm run check-format pass. The api-client tests also pass on Node 22, 24 and 26.

For reviewers

  • This applies to every client from createClient, including the CLI's interactive commands: a 429 there now waits up to 5 minutes, logged only under DEBUG=@argos-ci/api-client, instead of failing fast. A visible warning or a per-client setting could follow.
  • A Retry-After HTTP date is treated as invalid on purpose, since the API only sends seconds.

🤖 Generated with Claude Code

A 429 was returned as is, so one throttled request failed the whole
upload on sharded CI runs: "APIError: HTTP 429 Too Many Requests: Too
many requests, please try again later."

apiFetch now waits for Retry-After and retries, up to 5 times, waiting
at most 60s each time. The API counts requests in fixed 5-minute
windows, so this sees a request through a whole window while bounding
how long CI waits. A missing or invalid header falls back to exponential
backoff from minTimeout.

These retries are counted apart from the server error retries, keep the
request id and bump x-argos-retry-attempt, and aborting the request
stops the wait. Once they run out, the last 429 response is returned
unchanged, so callers report the same APIError.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@gregberge gregberge added the bug Something isn't working label Oct 2, 2026
@vercel

vercel Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
argos-js-sdk-reference Ready Ready Preview Oct 2, 2026 7:22am UTC

Request Review

@gregberge
gregberge requested a review from jsfez October 2, 2026 07:40
@gregberge
gregberge merged commit f5af9f4 into main Oct 2, 2026
77 checks passed
@gregberge
gregberge deleted the claude/modest-margulis-dbf31e branch October 2, 2026 08:55

This branch was successfully deployed

1 active deployment
Preview — 72a84a13 Deployed Oct 2, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant