Skip to content

MLX: release a temp slot once its chain is done with it - #23109

Merged
metascroy merged 3 commits into
pytorch:mainfrom
msluszniak:ms/mlx-release-temp-slots
Sep 26, 2026
Merged

metascroy merged 3 commits into
pytorch:mainfrom
msluszniak:ms/mlx-release-temp-slots

Conversation

@msluszniak

@msluszniak msluszniak commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

ExecutionState::tensors is never pruned, so a chain ends holding every intermediate it produced. Whether that costs anything depends on the AOT memory plan: whisper's reuses a handful of slots, but a gemma4-e2b prefill plan assigns roughly one slot per instruction, so its 3243 instructions finish with 1918 live tensors whose sizes scale with sequence length.

Measured on an iPhone 16, gemma4-e2b int4, 2392-character prompt:

live slots liveMB peak phys_footprint outcome
stock 1918 3102.9 MB 3401 MB killed by jetsam in prefill
this PR 3 1.2 MB 1379 MB completes normally

A 52-character prompt completes on both. Reported downstream as software-mansion/react-native-executorch#1489.

Only temps are eligible: ids run Constant -> Input -> Output -> MutableBuffer -> Temp, so >= mutable_buffer_end is neither a method output nor a caller-visible buffer. A chain containing SCAN or IF opts out, since for_each_tid does not descend into nested chains.

Independent of #22932's eval threshold: with the threshold off the intermediates stay lazy and the growth reappears in the final async_eval; with it on they are materialized and retained. This release is what makes the bounded path bounded.

cc @nil-is-all @metascroy

@pytorch-bot

pytorch-bot Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23109

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 64 Pending

As of commit 2bcb0d6 with merge base fa61c2b (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Sep 24, 2026
@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@nil-is-all nil-is-all added the module: mlx Issues related to MLX Backend: Metal-accelerated inference on Apple Silicon label Sep 24, 2026
@executorch-triage executorch-triage Bot added the community: contribution PRs coming from community (excluding hardware partners) label Sep 24, 2026
@msluszniak
msluszniak force-pushed the ms/mlx-release-temp-slots branch from 77b197c to 585f206 Compare September 24, 2026 19:18
@msluszniak
msluszniak force-pushed the ms/mlx-release-temp-slots branch from 585f206 to 8b77188 Compare September 24, 2026 19:24
@metascroy

metascroy commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

@msluszniak I'm wondering if these kinds of MLX memory controls would be better to expose than this eval setting we've added:

https://ml-explore.github.io/mlx/build/html/python/memory_management.html

What are your thoughts? Have you tried using these to see if they help on iOS? We could expose them via backend options like we do with clear cache.

@metascroy

Copy link
Copy Markdown
Contributor

@claude review this code?

@claude

claude Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Claude finished @metascroy's task in 2m 31s —— View job


Review: MLX temp-slot release

  • Gather context / read diff
  • Review MLXInterpreter.h release logic
  • Review tests
  • Post review

Verdict: I think the approach is correct and the change is well scoped. I found no correctness bugs. My main suggestion is to compute the liveness table once at load, not on every run_chain call. The rest are smaller points.

Correctness: checks that passed

  • Only temps are touched. id.idx >= mutable_buffer_end matches the Constant -> Input -> Output -> MutableBuffer -> Temp layout in ExecutionState::bind. output_map only points at outputs, so nothing execute() reads after run() gets dropped.
  • for_each_tid covers every Tid field kind. It handles tid, optional_tid, list_tid, vid_or_tid and int_or_vid_or_tid (serialization/generate.py:1026). Until now for_each_tid only fed the byte estimate, so a missed field would have been harmless. With this PR it decides when a slot is freed, so a missed field would free a live tensor. Anyone adding a new Tid-bearing field kind to the generator will now need to update the emitter too. A one-line note at _emit_cpp_for_each_tid would help.
  • SCAN and IF bodies are safe. ScanNode's sliced, outputs and carry, and anything the parent reads after an IF, are all named by the parent's SCAN/IF instruction. That marks them shared inside the body chain, so a carry that persists across iterations is never dropped. The parent chain opts out on its own because it contains the SCAN/IF.
  • Slot reuse from the AOT planner is safe. last_use is the slot's final reference across all its reuses, so reuse only delays the release.

Main suggestion: build last_use once at load, not per run_chain

compute_temp_last_use (MLXInterpreter.h:2088) walks every instruction in every chain and allocates two vectors sized to st.tensors. It does this on every call:

  • Every execute(): once per decode token. That is small next to GPU time, but it is fixed overhead on a hot path for data that never changes.
  • Every SCAN iteration: exec_scan calls run_chain on the body T times, so this costs O(T × total_instructions) per SCAN.

MLXProgram is immutable after load. A per-chain std::vector<std::vector<uint32_t>> (empty means opted out) could be built in MLXBackend init next to state.bind, then passed in or looked up by run_chain. This also keeps run() const without a mutable cache.

Fix this →

Smaller points

  1. SCAN/IF opt-out is all-or-nothing. Any main chain containing an IF or SCAN gets no release at all. A cheap improvement would count every tid a nested chain names (walked recursively) as a use at the SCAN/IF instruction's index in the parent. That is still conservative, and it would let those chains benefit too. This is fine as a follow-up.
  2. Dead bounds checks. tensor_index already throws out_of_range when idx >= tensors.size(), so the slot < shared.size() and slot < last_use.size() checks never fire. The has_value() check before reset() is also redundant, since reset() on an empty optional is a no-op.
  3. One comment and one test describe a case that can't happen in the backend. execute() calls h->state.reset() before every run() (MLXBackend.cpp:479), which clears all temps.
    • In the backend, a temp written by the init chain is already gone before the main chain runs. So "a temp the init chain produces and the main chain consumes is still live" (MLXInterpreter.h:2078) never happens.
    • The same reset() is why "its first mention must be a write" holds. That guarantee comes from reset(), not from this change.
    • TempSharedWithAnotherChainIsKept only passes because it calls interp.run without reset(). The guard is still worth keeping; the IF/SCAN-body case is the one it actually protects. Consider rewording the comment, and checking whether the Qwen 3.5 MoE failure mentioned in that test really came from a branch or body chain.
  4. Comment length. The ~30-line block above compute_temp_last_use repeats the PR description, including the gemma4 numbers. CLAUDE.md asks for minimal comments. Two or three lines would cover it: temps only, not named by any other chain, opt out when a SCAN/IF is present. The measurements belong in the commit message, where they already are.
  5. Missing test for a SCAN body. A small case where the body reads a carry or sliced tid it never writes would lock in the protection from the shared check. The IF case is already tested.
  6. mlx_eval_threshold_test rewrite. Switching to a probe output looks fine. EnabledAccountsEveryInstruction now checks one mid-chain value instead of all of them, which is expected now that temps are dropped.

On @metascroy's question about MLX memory controls

They look complementary to this change, not a replacement. set_memory_limit, set_cache_limit and clear_cache only act on memory MLX already considers free: the buffer cache and the allocation limit. The 1918 intermediates here are still referenced from ExecutionState::tensors, so MLX counts them as live and can't free them. A cache limit would cut down retained freed buffers once this release is in place. Exposing it through backend options like clear_cache_interval seems reasonable, but the growth described in this PR would still happen with it alone.

I have not built or run the tests. The C++ tests need Apple Silicon (Metal), and this runner is Linux.

@msluszniak
msluszniak force-pushed the ms/mlx-release-temp-slots branch from 8b77188 to efded24 Compare September 25, 2026 10:10
@msluszniak

msluszniak commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor Author

@msluszniak I'm wondering if these kinds of MLX memory controls would be better to expose than this eval setting we've added:

https://ml-explore.github.io/mlx/build/html/python/memory_management.html

What are your thoughts? Have you tried using these to see if they help on iOS? We could expose them via backend options like we do with clear cache.

From what I'm ready these options are valid for memory that is already considered freed.

Applied claude review.

@metascroy

Copy link
Copy Markdown
Contributor

A few follow-ups:

  1. Account for bytes before releasing last-use temps. Currently, release_temp_slots() runs before accumulate_instruction_bytes(), so the accounting skips released inputs. For shrinking operations, this undercounts the largest tensor touched and can miss an evaluation barrier. The order should be: account bytes, release dead temps, then check the threshold and evaluate the remaining live roots.

  2. Add a FULL → SUM regression test. Create a 4,096-element float32 temp, reduce it to a scalar, and use a 24,576-byte threshold. Verify that evaluation is triggered, the temp slot is released, and the result is correct. Existing equal-sized ADD chains don’t catch this accounting regression.

  3. Compute last-use information once at load time. Cache it per chain after the program and tensor-slot layout are finalized, preserving the IF/SCAN opt-out and cross-chain sharing protection. Rebuilding it in every run_chain() repeats allocations and whole-program scans on each decode step and SCAN iteration.

Separately, could you benchmark temp release with MLX’s native memory/cache limits against this custom eval threshold stuff. I want to make sure we aren't adding infra that could be done by just adding more control that MLX already offers.

`ExecutionState::tensors` is never pruned, so a chain ends holding every
intermediate it produced. Whether that matters is decided by the AOT memory
plan: whisper's reuses a handful of slots, but a gemma4-e2b prefill plan
assigns roughly one slot per instruction, so its 3243 instructions finish with
1918 live tensors. Their sizes scale with sequence length, so on a phone a long
enough prompt is an OOM kill. Measured on an iPhone 16, peak phys_footprint
over a 2392-character prompt drops from 3401 MB (killed) to 1379 MB.

Only temps are dropped, and only those no other chain names: a branch or scan
body is run through run_chain of its own, so a temp it writes for its caller,
or a carry it only reads, would otherwise die at its last use inside the body.
A chain holding a SCAN or IF opts out entirely, since for_each_tid does not
walk nested chains.

The tables are built once in ExecutionState::bind rather than per run_chain
call, which would otherwise walk every chain on each execute() and on each
scan iteration.
run_chain released last-use temps before charging the instruction, so a
shrinking op (FULL -> SUM) was charged for its scalar output instead of the
input it had just consumed, and could miss a barrier. Order is now: account,
release, then check the threshold. Covered by a FULL -> SUM regression test.

compute_temp_last_use also finds cross-chain temps in one pass instead of
walking every other chain per chain.
@msluszniak
msluszniak force-pushed the ms/mlx-release-temp-slots branch from efded24 to 75e79e4 Compare September 25, 2026 18:22
@metascroy

Copy link
Copy Markdown
Contributor

whisper's reuses a handful of slots, but a gemma4-e2b prefill plan assigns roughly one slot per instruction, so its 3243 instructions finish with 1918 live tensors

Out of curiosity, was this on a recent gemma4 export (within last month)? MLX now aggressively reuses temp slots during export, whereas before it didn't.

@msluszniak

Copy link
Copy Markdown
Contributor Author

Addressed in 75e79e4 (also rebased on main, which fixes the mypy failure):

  1. Bytes are accounted before release; the threshold check comes after.
  2. Added ShrinkingOpIsAccountedBeforeRelease (FULL 4096 f32 -> SUM, 24,576 B threshold). It fails on the old order.
  3. The tables were already built once in bind() at init; the cross-chain check is now a single pass.

Benchmark: M3 Pro, Qwen3-0.6B int4 exported from this branch, 2047-token prefill, temp release on, medians of 4 interleaved rounds x 5 runs. Peak is MLX peak over post-load active memory.

arm ms peak MB
none 632 829
eval threshold 512 MB 640 233
eval threshold 128 MB 654 97
set_memory_limit +512 MB 635 519
set_memory_limit +256 MB 635 263
set_memory_limit +128 MB 634 135
set_cache_limit(0) 638 829

Outputs identical; 1-token decode is 6.3 ms in every arm except set_cache_limit(0) (7.1).

So set_memory_limit bounds peak about as well as the threshold, without the sync cost, and the cache limit does not bound it. Temp release is still needed separately: this Qwen3 plan reuses slots so it is a no-op here, but gemma4's prefill ends with ~1900 live slots, which an allocator limit cannot free.

@msluszniak

Copy link
Copy Markdown
Contributor Author

whisper's reuses a handful of slots, but a gemma4-e2b prefill plan assigns roughly one slot per instruction, so its 3243 instructions finish with 1918 live tensors

Out of curiosity, was this on a recent gemma4 export (within last month)? MLX now aggressively reuses temp slots during export, whereas before it didn't.

No, it's older than that, so I can check the export with the latest changes that might have an impact.

@metascroy

Copy link
Copy Markdown
Contributor

@claude review this code

@claude

claude Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Claude finished @metascroy's task in 2m 50s —— View job


Review: MLX temp-slot release (after 75e79e4)

  • Gather context / read diff
  • Review runtime changes (MLXExecutor.h, MLXInterpreter.h, generate.py)
  • Review tests
  • Post review

Verdict: I found no correctness bugs. The three follow-ups from the last round are all addressed:

  • Bytes are now accounted before release.
  • There is a FULL → SUM regression test.
  • The last-use tables are built once in bind().

What's left below is mostly polish, plus one optional suggestion on the data structure.

Checks that passed

  • Only temps are released. output_map is built only from IdSpace::Output slots (program_builder.py:955). named_slots only feed buffer and constant loading. So nothing that execute() or the runner reads after run() can be a released temp.
  • Ordering in run_chain (MLXInterpreter.h:2041-2061). The order is now: account, release, ++idx, then check the threshold and evaluate. That is what @metascroy asked for. evaluate_state_tensors then only roots the live slots. Every released temp's consumers have already been dispatched, so nothing downstream loses its graph.
  • ShrinkingOpIsAccountedBeforeRelease really tests the ordering.
    • With the new order, SUM's widest tensor is the 16 KB temp0, so pending is 32 KB, which crosses the 24 KB threshold.
    • With the old order, temp0 is already gone, so SUM only adds 4 B.
  • for_each_tid covers every field kind. _get_field_kind raises on any type it does not recognise (generate.py:925,960). A new FBS type therefore can't slip past classification without someone noticing.
  • Cross-chain sharing is correct.
    • Anything a SCAN or IF names (sliced, outputs, carry, or a branch output the parent reads) ends up kShared.
    • A slot named only by an IF's then and else chains is also kShared, so it is kept. That is conservative but safe.
    • A body-local temp that is read before it is written would already throw uninitialized tensor on the first iteration. Release can't turn a working program into a silently wrong one here.
  • bind() happens before the init chain runs (MLXBackend.cpp:387), so the init chain gets its own table and releases its own temps. The table is read-only after bind(), so sharing it with multi-session rebinds is fine.

Suggestions

  1. Optional: store a release list per instruction instead of a slot → last-use map (MLXExecutor.h:184, MLXInterpreter.h:2071). The current layout has two costs:

    • Memory: every non-opted-out chain gets a tensors.size() vector, so memory is num_chains × num_slots. Programs with many small IF/SCAN chains pay for full-width tables that are almost entirely kNoTempLastUse.
    • Runtime: release_temp_slots re-runs for_each_tid plus tensor_index (with its range checks) on every instruction.

    A flattened table with offsets per instruction, holding the slots to drop, fixes both. It sizes memory to the number of actual releases, and the hot loop becomes for (slot : release_at[idx]) st.tensors[slot].reset();. This isn't a blocker, since the current cost is small next to the GPU work.
    Fix this →

  2. Enforce the generate.py note in code instead of a comment (generate.py:1026-1040). _emit_cpp_for_each_tid ends in return []. If someone adds a new Tid-bearing kind to _get_field_kind but forgets this emitter, it silently frees live tensors. Listing the known non-Tid kinds explicitly and raising on anything else would turn the comment into a check.
    Fix this →

  3. ScanBodyKeepsACarryItOnlyReads only checks the table (mlx_temp_release_test.cpp:226). It never runs the SCAN, because exec_scan needs originals[0] to get T. Adding one input as an original, with sliced as a temp, would let the test run interp.run(...) and check that the carry survives several iterations and the output is correct. The current test is still a reasonable guard.

  4. Comment trimming (CLAUDE.md asks for minimal comments).

    • "Computed once here rather than per run_chain, which would cost a walk of every chain on each execute() and each scan iteration" (MLXExecutor.h:179-181) records review history rather than describing the code, so it can go.
    • The four-line block at MLXInterpreter.h:2018-2021 repeats the table's doc comment. // Empty table: this chain opts out. is enough.
  5. TempSharedWithAnotherChainIsKept still exercises the init → main handoff. In the backend, state.reset() makes that handoff impossible, and the IF test covers the case that actually happens. Keep it or drop it; it's harmless.

  6. Nit, mlx_eval_threshold_test.cpp:43: std::numeric_limits is now used, so add #include <limits> rather than relying on a transitive include.

On the benchmark

Your numbers suggest set_memory_limit bounds peak memory about as well as the eval threshold, without the extra sync cost. That supports @metascroy's idea of exposing it as a backend option, and possibly dropping the custom threshold later. Temp release stays independent of that choice: an allocator limit can't free slots that ExecutionState still references. Re-measuring gemma4 with a current export, which now reuses temp slots aggressively, would show how much this PR still buys in practice. It stays a cheap safety net either way.

I have not built or run the tests. They need Apple Silicon with Metal, and this runner is Linux.

@metascroy

Copy link
Copy Markdown
Contributor

@claude review this code

@claude

claude Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Claude finished @metascroy's task in 3m 11s —— View job


Review: MLX temp-slot release (after 75e79e4)

  • Gather context / read diff
  • Review release logic in MLXExecutor.h / MLXInterpreter.h
  • Check which export paths emit IF/SCAN (these decide whether a chain opts out)
  • Review tests
  • Post review

Verdict: I found no correctness bugs. The three follow-ups from the last round are all fixed:

  • Bytes are now accounted before release, and the threshold check comes after.
  • The FULL→SUM regression test is added.
  • The tables are built once in bind().

The main remaining issue is how many models actually benefit from the release.

Correctness: checks that passed

  • Nothing the caller reads is freed. output_map is built only from IdSpace.Output (builder/program_builder.py:955), so the >= mutable_buffer_end filter cannot free a method output.

  • Cross-chain sharing is conservative enough. Take a tid that more than one chain names, for example:

    • a SCAN's sliced, outputs or carry;
    • a temp an IF branch writes and the parent reads.

    Any such tid becomes kShared, so no chain ever frees it. The gated_delta_rule body shows this working (custom_kernel_ops/gated_delta_rule.py:491-506):

    • carry is seeded by an IdCopy in the parent, so it is shared.
    • t0/t1/t2 are named only in the body, so the body owns them. Each iteration writes them before reading them, so freeing them at their last use inside the body is safe.
  • The init chain doesn't reduce what gets freed. emit_init is only used for constant precompute today (op_helpers.py:457), so no temp ends up shared between init and main.

  • The order in run_chain is right: dispatch → account → release → threshold eval. The barrier no longer roots dead temps.

Main issue: IF/SCAN opt-out skips the models that need this most

The main chain opts out entirely if it contains any IF or SCAN. At export, IF is emitted whenever the condition is a runtime value. That covers:

  • Every GGUF q4k/q5k/q6k linear (custom_kernel_ops/gguf/q{4,5,6}k/linear.py, condition M-1, where M is the number of activation rows). Any dynamic-seq-len GGUF model gets no release at all.
  • MoE sorted/unsorted dispatch (ops.py:2233, ops.py:2300).
  • On-device sampling (ops.py:4848, temperature → greedy/sample).
  • SCAN from gated_delta_rule (Qwen 3.5 style linear attention).

These are the long-prefill models where retained temps hurt most. On such a model, the gemma4 number would silently fall back to the stock behavior.

I'd suggest a fix with the same safety properties. When building a parent chain's table, treat every tid named anywhere inside an IF's or SCAN's nested chains (walked recursively, as init_chain_references_mutable_buffer already does) as used at the IF/SCAN instruction's index. Then:

  • The parent can free those tids after the IF/SCAN returns.
  • Nested chains keep the current exclusive-owner rule, so bodies and branches still never free anything their caller touches.

A slot is then shared only between the parent and its own nested chains, not across unrelated chains. The IF test would need its expectation flipped: kShared becomes releasable after out = shared + shared.

Fix this →

Smaller points

  1. Table size and how it is stored. temp_last_use is dense: num_chains × tensors.size() uint32_ts. Each GGUF linear adds two branch chains. So a 28-layer model with 7 linears per layer has about 393 chains, and with a few thousand temps that is several MB of mostly kNoTempLastUse.

    A sparse table would fix this: per chain, per instruction, the list of slots to drop (std::vector<std::vector<std::vector<uint32_t>>>, or flattened with offsets). Its size tracks the total number of tid references. It would also let release_temp_slots skip the for_each_tid walk and the tensor_index calls in the hot loop.

  2. Load-time behavior change. bind() now calls for_each_tid on every instruction of every chain. for_each_tid's default: case throws (MLXLoader.h.tmpl:134), and before this PR that path only ran with the threshold on. The generator covers every node table, so this is fine today. It does mean a gap in the generator now fails program load, not just the eval-threshold path. That seems like the right way to fail, but it deserves a line in the commit message.

  3. TempSharedWithAnotherChainIsKept (mlx_temp_release_test.cpp:149) is still labelled as testing the init chain. In the backend, execute() calls reset() first, so the case it describes can't happen there. It still exercises the owner check, but a label like "chain 1 (another chain)" would avoid implying that init→main temp handoff is supported.

  4. ScanBodyKeepsACarryItOnlyReads never executes the SCAN; it only inspects the table. Its ScanNode also has no originals, so running it would crash. That is fine for a table test. Either rename it, or reword the comment so it doesn't claim the carry "stays live across iterations".

On the benchmark / MLX memory controls

These numbers point one way: set_memory_limit bounds peak memory about as well as the eval threshold, and it doesn't add the sync cost. That is a reasonable argument for exposing the limit as a backend option and possibly retiring the threshold later.

Two caveats before swapping:

  • set_memory_limit is process-global in MLX, not per-handle. A per-method backend option would leak across every loaded model in the process.
  • The limit only reclaims memory that isn't referenced. So it doesn't replace this PR: without the release, the ~1900 live gemma4 slots are still referenced.

It would help to repeat the benchmark on a model where release actually changes the live-slot count. Qwen3-0.6B's plan already reuses slots, so release has no effect on it. A fresh gemma4 export would do, if the newer slot reuse still leaves a large live set.

I have not built or run the tests. They need Apple Silicon (Metal), and this runner is Linux.

@metascroy metascroy left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM! Just running tests before merging.

@metascroy

Copy link
Copy Markdown
Contributor

This PR looks good to improve memory on previously exported pte files, but I still wonder if it still does much with the latest export side changes.

On the new eval_memory_threshold option: do you think it's worth keeping around if we add an option for set_memory_limit?

@msluszniak

msluszniak commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor Author

Yeah, my primarily target was 1.4.1 version I'm still using so the fact that the problem is resolved on newer version changes that a bit. We maybe drop this part to be honest and I will keep it till bump to ET 1.5.1 locally.

Regarding set_memory_limit, i think set_memory_limit would be sufficient if added, at least benchmarks suggest so. The only thing that I'm unsure is that set_memory_limit is process-global and needs an absolute value, so might not be flexible with scenarios with multiple models.

@pytorch-bot

pytorch-bot Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Workflows were awaiting approval. CI has now been triggered for the ciflow labels on this PR.

@metascroy
metascroy merged commit 4199643 into pytorch:main Sep 26, 2026
283 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/mlx CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. community: contribution PRs coming from community (excluding hardware partners) module: mlx Issues related to MLX Backend: Metal-accelerated inference on Apple Silicon

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants