Skip to content

[DE-8270] Model weights upload & download (SDK side) - #469

Open
luke-e-schaefer wants to merge 3 commits into
masterfrom
lukeschaefer/de-8270-upload-download-model-weights
Open

[DE-8270] Model weights upload & download (SDK side)#469
luke-e-schaefer wants to merge 3 commits into
masterfrom
lukeschaefer/de-8270-upload-download-model-weights

Conversation

@luke-e-schaefer

@luke-e-schaefer luke-e-schaefer commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

Summary

SDK side of model weights upload / download, mirroring the REST surface merged in scaleapi#149063.

Stacked on #467 — base is update-nuc-sdk-for-new-eval-stuff-pt1, so review only the top commit. Retarget to master once 467 lands.

The two primary methods are the ones the server PR's API-docs pages (ApiDocsPage/models-python/{upload,download}-model-weights.md) already document, so the published docs and the SDK agree:

import nucleus

client = nucleus.NucleusClient(YOUR_SCALE_API_KEY)
model = client.get_model(reference_id="My-CNN")

client.upload_model_weights(model, "/path/to/weights.bin")
client.download_model_weights(model, "/path/to/save/weights.bin")

Added

  • NucleusClient.upload_model_weights(model, path, *, content_type=None, original_filename=None, checksum_sha256=None, on_progress=None) — presign → PUT direct to storage → finalize. Returns ModelWeights.
  • NucleusClient.download_model_weights(model, path, *, on_progress=None) — resolves the signed URL and streams to disk (creating parent dirs). Returns the path written.
  • get_model_weights(model) / delete_model_weights(model) for the remaining two routes.
  • Model.upload_weights() / .download_weights() / .weights() / .delete_weights() — thin delegation to the client, matching how Benchmark wraps its client methods.
  • ModelWeights metadata type: present, status, size_bytes, original_filename, content_type, download_url.

All model arguments accept either a Model or a bare model id (prj_*).

Notes for review

  • Bytes never transit the Nucleus API. Transfers go straight to storage over presigned URLs, so multi-GB artifacts aren't subject to API request-size limits.
  • Part PUTs are sent with no headers. They're signed without the Content-Type condition, so forwarding requiredHeaders (which the single PUT does need) makes S3 reject the signature. There's a test pinning this in both directions — it's the easiest thing to get wrong here, and the frontend hook in 149063 has the same split.
  • Download uses ?json=1 to fetch the signed URL rather than following the 302, so the API's auth headers are never sent to storage.
  • Multipart above 5 GB, 4 parts in flight (a single S3 PUT is connection-throughput-bound). Missing part ETags fail on the first part rather than after transferring everything and dying at finalize.
  • The 10 GB server cap is checked client-side before presign, so an oversized file fails without a network round-trip.
  • The weights routes serialize camelCase in both directions, unlike most endpoints this SDK talks to, so the new payload keys are grouped and labelled as such in constants.py.

Tests / Version

  • tests/test_model_weights.py — 24 mock-based unit tests (no live API, no real S3): DTO parsing, payload builders, single vs. multipart transfer, header split, ETag/failure handling, progress callbacks, download streaming, all four client methods, and the Model wrappers.
  • Verified locally the way CI does: pylint nucleus 10.00/10, mypy --ignore-missing-imports nucleus clean, ruff clean, black + isort clean, 61 mock-based tests passing across the eval/benchmark/leaderboard/weights suites.
  • pyproject.toml0.19.1 + CHANGELOG entry (patch bump: additive new methods, per CLAUDE.md).

resolves https://linear.app/scale-epd/issue/DE-8270

🤖 Generated with Claude Code

Greptile Summary

This PR adds model weights upload and download to the Nucleus Python SDK. Bytes flow directly between the caller and storage over presigned URLs (presign → PUT → finalize), so multi-GB artifacts never transit the Nucleus API.

  • NucleusClient gains upload_model_weights, download_model_weights, get_model_weights, and delete_model_weights; Model gets matching thin wrappers (upload_weights, download_weights, weights, delete_weights).
  • Single-part uploads (< 5 GB) use a _ProgressReader wrapper to deliver incremental progress callbacks; multipart uploads (≥ 5 GB) run up to 4 concurrent part PUTs with a threading.Lock-protected counter, correctly sending no headers on part PUTs to avoid S3 signature rejection.
  • The new ModelWeights dataclass represents artifact metadata; 24 mock-based unit tests cover all paths including the header-split invariant, ETag failure handling, and progress callbacks.

Confidence Score: 5/5

  • This PR is safe to merge. The upload/download logic is well-structured, the thread-safety concern on the multipart progress counter is already addressed with a lock, and the 24-test suite covers the critical edge cases.
  • The change is additive, bytes never transit the Nucleus API, and the presign→PUT→finalize protocol correctly handles both the single-part and multipart code paths including the header-split invariant. No existing behaviour is modified.
  • nucleus/model_weights.py — the _client field in ModelWeights participates in __eq__, which is a minor design issue but not a correctness problem for any current caller.

Important Files Changed

Filename Overview
nucleus/model_weights.py New module implementing presign→PUT→finalize upload (single and multipart) and streamed download. Well-structured with proper thread locking on the transferred counter. The ModelWeights._client field participates in dataclass __eq__ by default, which can cause unexpected inequality between logically identical objects.
nucleus/init.py Adds four new NucleusClient methods and imports private helpers from model_weights (needed because the methods live in init.py and tests patch at the nucleus.* level). ModelWeights is correctly added to all.
nucleus/model.py Adds four thin delegation methods (upload_weights, download_weights, weights, delete_weights) mirroring the Benchmark pattern. Clean and correct.
nucleus/constants.py Adds 16 camelCase constants for the weights wire format. Well-grouped and labelled to explain the deviation from the rest of the file's snake_case style.
tests/test_model_weights.py 24 mock-based unit tests covering DTO parsing, payload builders, single/multipart upload (including the header-split invariant), progress callbacks, download streaming, all four client methods, and the Model wrappers. Test coverage is comprehensive and well-organised.

Sequence Diagram

sequenceDiagram
    participant User
    participant SDK as NucleusClient
    participant API as Nucleus API
    participant S3 as Storage (S3)

    Note over User,S3: Upload flow
    User->>SDK: upload_model_weights(model, path)
    SDK->>SDK: check file size ≤ 10 GB
    SDK->>API: "POST model/{id}/weights/presign"
    API-->>SDK: presign response

    alt Single PUT (uploadUrl present)
        SDK->>S3: PUT presignedUrl (with requiredHeaders)
        S3-->>SDK: ETag
    else Multipart (parts[] present)
        par Up to 4 concurrent part uploads
            SDK->>S3: PUT part[1].url (no headers)
            SDK->>S3: PUT part[2].url (no headers)
            SDK->>S3: PUT part[N].url (no headers)
        end
        S3-->>SDK: ETags per part
    end

    SDK->>API: "POST model/{id}/weights/finalize"
    API-->>SDK: ModelWeights DTO
    SDK-->>User: ModelWeights

    Note over User,S3: Download flow
    User->>SDK: download_model_weights(model, path)
    SDK->>API: "GET model/{id}/weights/download?json=1"
    API-->>SDK: "{"url": signedUrl}"
    SDK->>S3: GET signedUrl (streaming)
    S3-->>SDK: file bytes (chunked)
    SDK-->>User: path written
Loading

Reviews (5): Last reviewed commit: "fix(weights): lock the multipart progres..." | Re-trigger Greptile

Comment thread nucleus/model_weights.py
Comment thread nucleus/model_weights.py
Base automatically changed from update-nuc-sdk-for-new-eval-stuff-pt1 to master August 11, 2026 14:24
luke-e-schaefer and others added 2 commits August 11, 2026 14:31
Mirrors the REST surface shipped in scaleapi#149063 so users can attach a
weights artifact to a model and fetch it back from Python.

- `NucleusClient.upload_model_weights` / `download_model_weights` — the two
  methods the server PR's API-docs pages already document — plus
  `get_model_weights` and `delete_model_weights` for the remaining routes.
- `Model.upload_weights()` / `download_weights()` / `weights()` /
  `delete_weights()` delegate to the client, matching how `Benchmark` does it.
- New `ModelWeights` metadata type parsed from the weights DTO. The weights
  routes serialize camelCase both ways, unlike most of this SDK's endpoints,
  so the new payload keys are grouped and labelled in `constants.py`.
- Transfers go straight to storage via presigned URLs and never through the
  API, so artifacts aren't subject to API request-size limits. Over 5 GB the
  server hands back multipart parts, which upload 4 at a time; `on_progress`
  reports `(bytes_transferred, total_bytes)`.
- Size is checked against the server's 10 GB cap before presign, so an
  oversized file fails without a network round-trip.

Two things worth knowing for review: part PUTs must be sent with *no* headers
(they're signed without the Content-Type condition, so forwarding
`requiredHeaders` makes S3 reject the signature), and download resolves the
signed URL via `?json=1` rather than following the 302, so the API's auth
headers are never sent to storage.

24 mock-based unit tests in `tests/test_model_weights.py`; version bumped to
0.19.1 (additive, per CLAUDE.md).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The docstrings are what users read, so they shouldn't describe how the
artifact gets stored or moved. Dropped the presign/multipart/direct-to-storage
narration from every public docstring, the `ModelWeights` attribute docs, and
the CHANGELOG, leaving what a caller actually needs: what the method does, who
can call it, the size limit, and the arguments.

Also made the transfer helpers private (`_presign_payload`,
`_transfer_weights_to_storage`, `_stream_weights_to_file`,
`_finalize_payload`) so the mechanics don't show up in the generated API docs
at all, rather than only being reworded.

Kept the two in-body comments that explain why part uploads send no headers
and why the download URL is fetched as JSON — those aren't user-visible and
each one guards a real footgun.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@luke-e-schaefer
luke-e-schaefer force-pushed the lukeschaefer/de-8270-upload-download-model-weights branch from df6ede7 to 83c25f7 Compare August 11, 2026 14:35
…progress

Addresses two review comments:
- transferred += len(chunk) ran unsynchronized across the part-upload pool,
  so concurrent workers could drop updates. The counter and the value handed
  to on_progress are now taken under a lock.
- A single PUT reported nothing until it finished, then jumped to 100%.
  When a callback is supplied the body is wrapped so progress comes from the
  read side; the wrapper delegates everything but read(), so requests still
  sizes the body from fileno()/tell() and sends Content-Length as before.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant