Skip to content

feat(upload): split profile archive uploads into concurrent parts - #565

Open
lvaroqui wants to merge 1 commit into
mainfrom
cod-3700-multipart-upload
Open

lvaroqui wants to merge 1 commit into
mainfrom
cod-3700-multipart-upload

Conversation

@lvaroqui

@lvaroqui lvaroqui commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Upload profile archives of 64 MiB and more as S3 multipart uploads, with parts sent concurrently.

S3 rejects a single upload request above 5 GiB, so large walltime and memory profile archives could not be uploaded (walltime folders above 5 GiB were gzipped on disk to try to fit). Even below that limit, a single connection to S3 only reaches about 20-25 MiB/s on GitHub-hosted runners, so a 1 GiB archive took close to a minute to upload.

How it works

  • Part layout: archives of 64 MiB and more are split in about two parts per concurrent upload, each between 16 MiB and 256 MiB, so a slow part doesn't leave the other upload slots idle at the end. The md5 of every part is computed in the same pass over the archive as the archive md5, and sent in the upload metadata (bumped to version 12) as profileMultipart (size, partSize, partMd5s).
  • Upload: the upload endpoint answers such a request with multipartUploadUrls (partUrls, completeUrl) instead of uploadUrl. Each part URL is presigned with the part's Content-MD5, so S3 checks every part. Parts are uploaded 8 at a time, each with its own retries, then the ETags are sent in part order to the completion URL. S3 can report a failed completion with a 200 status, so the response body is also checked for an error.
  • Both archive kinds: on-disk (walltime, memory) archives are streamed from disk, and in-memory gzip (simulation) archives are sliced without copying.
  • S3 specifics (request headers, ETag, completion body) live in the new upload::s3 module.
  • CODSPEED_UPLOAD_CONCURRENCY overrides the number of concurrent part uploads.

Archives below 64 MiB keep the single upload request, and their metadata is unchanged apart from the version.

Other changes

  • Walltime profile folders above 5 GiB are no longer gzipped on disk, and the maximum archive size goes from 5 GiB to 15 GiB.
  • Archives are hashed while streaming on the blocking thread pool, instead of being read whole into memory on the async runtime.
  • Single-request uploads of in-memory archives now go through the same retry loop as on-disk ones.

Measurements

Concurrency sweep on a 6 GiB walltime archive (25 parts of 256 MiB), which led to the default of 8:

Concurrency ubuntu-latest CodSpeed macro runner (Ryzen 9950X)
1 240.3s · 25.6 MiB/s 59.5s · 103.2 MiB/s
2 83.4s · 73.7 MiB/s 33.9s · 181.4 MiB/s
4 57.2s · 107.4 MiB/s 21.3s · 289.0 MiB/s
8 45.2s · 135.8 MiB/s 8.2s · 746.5 MiB/s
16 48.7s · 126.1 MiB/s 10.3s · 596.5 MiB/s

Smaller archives with the final part layout, concurrency 1 (close to the previous single request) vs 8:

Archive Runner Concurrency 1 Concurrency 8
256 MiB ubuntu-latest 11.6s 2.8s (4.1×)
1 GiB ubuntu-latest 55.1s 9.0s (6.1×)
256 MiB macro runner 2.7s 1.3s (2.1×)
1 GiB macro runner 10.1s 1.6s (6.3×)

The backend support for profileMultipart is not released yet, so the multipart path cannot be verified end to end against production for now.

Closes COD-3700

@codspeed

codspeed Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

✅ 31 untouched benchmarks
⏩ 6 skipped benchmarks1


Comparing cod-3700-multipart-upload (b2177c0) with main (8c603c2)

Open in CodSpeed

Footnotes

  1. 6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@lvaroqui
lvaroqui force-pushed the cod-3700-multipart-upload branch from 454fdc6 to 1cc552d Compare October 6, 2026 13:30
@lvaroqui lvaroqui changed the title feat(upload): upload profile archives over 5 GiB in parts feat(upload): upload profile archives using S3 multipart in parts Oct 6, 2026
@lvaroqui
lvaroqui force-pushed the cod-3700-multipart-upload branch from 1cc552d to a2f09f3 Compare October 6, 2026 13:34
@lvaroqui lvaroqui changed the title feat(upload): upload profile archives using S3 multipart in parts feat(upload): split profile archive uploads into concurrent parts Oct 6, 2026
@lvaroqui
lvaroqui requested review from GuillaumeLagrange and removed request for GuillaumeLagrange October 6, 2026 13:37
@lvaroqui
lvaroqui marked this pull request as ready for review October 6, 2026 14:17
@greptile-apps

greptile-apps Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[High risk] Changes how profile archives are uploaded to cloud storage.

The PR does not appear safe to merge until retryable completion errors outside the current allowlist are handled.

Fix All in Claude CodeFindings

  1. P1 Retryable completion errors fail ▶
Fix with agent prompt
### Issue 1
src/upload/s3.rs:95-99
If S3 returns a retryable error such as `Throttling` inside a 200 completion response, this four-code list treats it as permanent. The caller then stops retrying, so an upload that could succeed on a later completion attempt fails.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR splits large profile archives into concurrently uploaded S3 parts, adds multipart metadata and completion handling, and retains single-request uploads for smaller archives. The latest revision adds bounded reads for non-success error bodies and distinguishes transient from permanent completion errors.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Archive and hash] --> B{At least 64 MiB?}
  B -- No --> C[Single S3 upload]
  B -- Yes --> D[Concurrent part uploads]
  D --> E[Collect ordered ETags]
  E --> F[Complete multipart upload]
  F --> G{Completion response}
  G -- Success --> H[Done]
  G -- Classified transient --> F
  G -- Classified permanent --> I[Fail]
Loading

Reviews (3) · Last reviewed commit: "feat(upload): split profile archive uplo..."

Comment thread src/upload/uploader.rs
Comment thread src/upload/s3.rs Outdated
Comment thread src/upload/interfaces.rs
@lvaroqui
lvaroqui force-pushed the cod-3700-multipart-upload branch from a2f09f3 to c60ccf8 Compare October 6, 2026 15:38
Comment thread src/upload/uploader.rs Outdated
Comment thread src/upload/uploader.rs Outdated
S3 rejects single uploads above 5 GiB, and a single connection to S3
only reaches about 20-25 MiB/s on GitHub-hosted runners, so large
profile archives were slow or impossible to upload.

Archives of 64 MiB and more are now sent as an S3 multipart upload, in
about two parts per concurrent upload, each between 16 MiB and 256 MiB.
The md5 of every part is computed in the same pass as the archive md5
and sent in the upload metadata (version 12) as `profileMultipart`. The
API then answers with `multipartUploadUrls`: the parts are uploaded 8 at
a time (overridable with `CODSPEED_UPLOAD_CONCURRENCY`), each with its
own retries, then the upload is completed with the part ETags in order.
This applies to both on-disk and in-memory (gzip) archives.

Walltime profile folders above 5 GiB are no longer gzipped on disk to
fit in a single request, and the maximum archive size goes from 5 GiB to
15 GiB. Archives are now hashed while streaming on the blocking thread
pool, instead of being read whole into memory on the async runtime.

Closes COD-3700
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lvaroqui
lvaroqui force-pushed the cod-3700-multipart-upload branch from c60ccf8 to b2177c0 Compare October 6, 2026 16:13
Comment thread src/upload/s3.rs
Comment on lines +95 to +99
pub(super) fn is_transient(&self) -> bool {
matches!(
self.code.as_deref(),
Some("InternalError" | "ServiceUnavailable" | "SlowDown" | "RequestTimeout")
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Retryable completion errors fail

If S3 returns a retryable error such as Throttling inside a 200 completion response, this four-code list treats it as permanent. The caller then stops retrying, so an upload that could succeed on a later completion attempt fails.

Prompt To Fix With AI
This is a comment left during a code review.
Path: src/upload/s3.rs
Line: 95-99

Comment:
**Retryable completion errors fail**

If S3 returns a retryable error such as `Throttling` inside a 200 completion response, this four-code list treats it as permanent. The caller then stops retrying, so an upload that could succeed on a later completion attempt fails.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code Fix in Codex

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You’re right — Throttling is not an S3 error code; S3 uses SlowDown for throttling, which this list already handles. I withdraw the comment.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant