Repository navigation
Releases: ml-explore/mlx
Releases · ml-explore/mlx
Release list
v0.32.3
What's Changed
- Fix scan and sort for a zero-size axis by @eyupcanakman in #4340
- Add a correction parameter in std and var by @prady0t in #4348
- Fix tests/run.py stuck in macOS CI by @zcbenz in #4410
- Extension Example Updated by @jagrit06 in #4412
- Fix sorted gather_qmm NAX row overflow above 32K by @PhilipJohnBasile in #3922
- Fix deadlock caused by mx.clear_streams() holding GIL by @aleroot in #4413
- Fix integer pow zeroing a whole SIMD vector on a negative exponent by @ayaangazali in #4354
- Make concurrency cap on load adaptative so that I/O scales with the machine by @aleroot in #4408
- Fix pad with an axes subset and negative axes by @kapellirohith in #4364
- Use sdpa_vector_2pass_1_gqa for GQA size 12 and 16 by @dudududukim in #4380
- python: Make axis of put_along_axis optional by @ayaangazali in #4360
- python: Make axis default to -1 in take_along_axis by @aaishwarymishra in #4368
- Propagate NaN in arg reductions by @atirna in #4291
- Fix mx.from_fp8 E4M3FN NaN decode with branchless carry by @saud5150 in #4376
- Update rules for PR limitation bypass list by @zcbenz in #4415
- [CUDA] Use native events for GPU fence waits by @strayberry in #4401
- Validate GGUF tensor dimensions by @roshaninfordham in #4378
- Fix crash on ellipsis indexing with too many trailing indices by @Adityaj0 in #4396
- [CUDA] Cholesky via cuSOLVER by @sashko-zakharchuk in #4208
- python: Fix setitem with negative index preceded by None by @Adityaj0 in #4414
- Update nanobind to 3.0.1 by @XXXXRT666 in #4417
- chore: Use normalize_axis_index in scan ops by @Adityaj0 in #4383
- python: Add explicit index and bytes support by @aaishwarymishra in #4388
- docs: fix typo recived -> received by @vaibhav8a in #4424
- Fix Metal grid sizing for strided scans by @TheDarkchip in #4420
- Fix SGD and Adafactor weight decay mutating caller arrays by @2sumtech in #4426
- Widen float16/bfloat16 to float32 in cpu reduce by @Ved235 in #4387
- sdpa_vector_2pass_1_gqa kernel batch offset for K/V in gqa decode by @dudududukim in #4431
- Make head-dim-256 prefill with array mask run fused NAX kernel by @dwijenpatel in #4416
- python: Make unstack return tuple instead of list by @JasonHonKL in #4448
- python: Make iter() throw TypeError for 0-dim array by @simeetnayan81 in #4425
- Fix save and save_safetensors corrupting a lazily loaded source file by @Cdimoy in #4434
- [Metal] Use hypot for complex ops by @JasonHonKL in #4345
- Add get_array_buffer_size to query buffer size for arrays by @aleroot in #4436
- Add M5 ultra tunings for non-quantized matmuls by @jagrit06 in #4447
- [CUDA] Fix completion worker busy loop by @strayberry in #4452
- Fix non transposed affine qmm dispatch logic by @RohanGautam in #4392
- Fix sorted gather_qmm on ragged K by @erwinzhang7 in #4009
- Don't restore a thread-affine stream when StreamContext dies on another thread by @michalk8 in #4462
- Fix Pad vjp with an axes subset and negative axes by @nileshpatil6 in #4441
- Fix integer div-by-zero crash in CPU backend (sort, scan, binary ops) by @MarcosAsh in #4442
- Use 64-bit file seeks on Windows by @dhiltgen in #4456
- Use precise::exp in Sigmoid so compiled and eager sigmoid agree by @pierre427 in #4461
- Use NAX attention for long unmasked D72/D80 inputs by @wyanzhao in #4455
- Add D512 support to Metal vector attention by @wyanzhao in #4459
- Add a scoped_env utility for tests by @zcbenz in #4474
- [Metal] global scale for qmm by @nastya236 in #4458
- Fix re-entering the same stream context manager by @michalk8 in #4478
- leak global CommandEncoder -- avoid cuda synchronize on process shutdown by @davidkoski in #4480
- Update CONTRIBUTING.md by @zcbenz in #4469
- python: fix bytes() on non-contiguous arrays by @axiom-of-choice in #4449
- Break the sibling cycle when an array is released by assignment by @tudalex in #4453
- Win: Use DXGI for accurate WDDM VRAM memory budget by @dhiltgen in #4457
- Add type hints and docstrings for ALiBi layer by @Ritabanm in #4472
- Fix fp quantized matmul corruption when the quantized dim is not a multiple of 32 by @kapellirohith in #3912
- Fix CUDA test synchronization flake by @dhiltgen in #4490
- Skip bypass list updates in forks by @XXXXRT666 in #4495
- Load global scales in qmm_t kernels by @dhiltgen in #4483
- Release the GIL when copying Metal DLPack inputs (mx - pytorch mps bridge) by @WindChimeRan in #4497
- Use matrix kernels for global-scale gather_qqmm by @dhiltgen in #4481
- Adding metal kernels for the gated delta nets. by @tpegolotti in #4020
- Fix jaccl ring all_gather: direction 1 slice was not mirrored by @Drifter4242 in #4443
- Skip Metal-only gated delta kernel tests on other backends(fix CI) by @aleroot in #4522
- [CUDA] Add global scale support to gather_qmm by @dhiltgen in #4507
- [BUG][Metal] Deadlock in fence by @nastya236 in #4552
- Fix ordering of complex64 with NaNs by @louen in #4519
- Improve metal memory usage for SDPA D256 by @dhiltgen in #4505
- Add support for matrix transpose by @aaishwarymishra in #4402
- Improve metal memory usage for SDPA D512 by @dhiltgen in #4487
- Support shapeless compilation of scan operations by @keeeeenw in #4510
- [Metal] Gather mm improvement by @nastya236 in #4567
- restoring perf regression for upsample
linearmode align_corners=False by @sp4s-s in #4500 - Add fused Metal kernels for fast.cross_entropy by @zsun6 in #4520
- Add missing defaults to tri, tril, triu, gather_mm signatures by @rishabhsai in #4492
- Add Metal SDPA support for D96/V64 by @wyanzhao in #4499
- Raise IndexError for out of bounds axes by @devangpratap in #4484
- Fix floor_divide for integers by @aaishwarymishra in #4515
- Fix cpu binary ops on large data by @Prudctual in #4517
- Do not throw from CUDA destructors and avoid implicit default streams in eval/compile by @aleroot in #4514
- [Metal] gather_qmm improvement by @nastya236 in #4572
- Fix floating-point constant precision in compiled kernels by @Ryan11c in #4511
- chore: Fix assertion in ReduceScatter::eval_cpu by @ronaldmannak in #4557
- Leak thread local streams on exit in main thread by @zcbenz in #4576
New Contributors
- @atirna made their first contribution in #4291
- @saud5150 made their first contribution in #4376
- @roshaninfordham made their first contribution in #4378
- @vaibhav8a made their first contribution in #4424
- @TheDarkchip made their first contribution in #4420
- @2sumtech made their first contribution in #4426
- @simeetnayan81 made their first contribution in https://github.com/ml-explore/...
v0.32.2
What's Changed
- Preserve subnormal float values when casting to bool by @reckylurker in #4224
- Support assigning through a bare Ellipsis index by @Adityaj0 in #4314
- Fix divmod truncating the quotient for floats by @ayaangazali in #4108
- Add force_fused option to scaled_dot_product_attention by @hojin12312 in #4185
- Reject a negative eps in the normalization layers by @ayaangazali in #4312
- Bound GGUF metadata string/array values against the file mapping by @x14ngch3n in #4212
- Read each K/V byte once in gqa-8 decode attention by @dudududukim in #4077
- Fix fft vmap and jvp for transforms over a subset of axes by @kapellirohith in #4138
- Fix median dropping NaN by @devteamaegis in #4146
- Fix the CPU scan over a size one axis with a padded stride by @kapellirohith in #4139
- Validate the optimizer betas at construction by @ayaangazali in #4310
RMSNormVJPbackward writes a full{n_rows, D}gw_tempintermediate by @JasonHonKL in #4293- [Bug]: add default none value to axis parameter of the take_along_axis by @aaishwarymishra in #4357
- Add a fused full-attention path for head_dim 256 on NAX devices by @wyanzhao in #3842
- Update nanobind to 2.15.0 by @XXXXRT666 in #4337
- Skip unnecessary simdgroup computations for quantised MOE matmuls on NAX by @RohanGautam in #4352
- Add AI usage policy by @zcbenz in #4331
- Raise cpu stream errors from synchronize by @robertomeroni in #4338
- chore: Validate eps in Adam at construction by @vraj00222 in #4361
- Bound winograd conv2d working set by tiling the batch by @Gusanidas in #4102
- Use a 32-row block in qmm_t_nax when one block covers all of M by @dwijenpatel in #4171
- Deduplicate fftshift and ifftshift by @Adityaj0 in #4318
- Fix Log and Equal is_equivalent ignoring primitive state by @kapellirohith in #4266
- Stabilize reduced-precision InstanceNorm by @ternaus in #4230
- Normalize negative axes in sort and argsort by @deBrian07 in #4332
- Clean up main thread compile cache before python interpreter shuts down by @zcbenz in #4373
- chore: Check malformed jaccl hostfile that miss rdma in pairs by @erwinzhang7 in #4284
- Round mxfp8 block scales up to avoid saturation by @dhiltgen in #4353
- Add support for the array_namespace_info by @aaishwarymishra in #4334
- Stop a failed CUDA graph commit from poisoning the encoder by @strayberry in #4356
- Fix quantized kernels in JIT build by @dwijenpatel in #4372
- [Metal][Performance] Avoid zero work in stride-2 ConvTranspose3d by @ternaus in #4343
- [CUDA] Ce fused kernel by @nastya236 in #3947
- Fix cpu exclusive scan for complex numbers by @ayaangazali in #4272
- Support Relocatable CUDA DLLs on Windows by @dhiltgen in #4382
- Use cast_to for fused AsType in compiled Metal kernels by @katlun-lgtm in #4351
- Declare DLPackCompatible protocol members as methods, not settable attributes by @Adityaj0 in #4384
- Fix quantizing sliced arrays by @zcbenz in #4381
- Fix einsum dropping a trailing empty subscript by @Adityaj0 in #4299
- Add script to run python tests by @zcbenz in #4393
- Hold GIL in AttachedData destructor by @zcbenz in #4391
New Contributors
- @hojin12312 made their first contribution in #4185
- @x14ngch3n made their first contribution in #4212
- @vraj00222 made their first contribution in #4361
- @ternaus made their first contribution in #4230
- @strayberry made their first contribution in #4356
Full Changelog: v0.32.1...v0.32.2
v0.32.1
What's Changed
- Fix int64 type cast error when loading GGUF metadata arrays by @danlee2002 in #3823
- Document default values in normalization layer docstrings by @Pablosinyores in #3819
- Warn at configure time when NAX kernels are disabled by @pierre427 in #3824
- [CUDA][Improvement] RMSNorm forward speed up by @nastya236 in #3850
- Fix captured random state in compile by @angeloskath in #3828
- Fix JIT preamble header filter matching project paths containing "Xcode" by @apocryphx in #3873
- Document default value of p in dropout layer docstrings by @ayaangazali in #3870
- [WIP] [CUDA] fsdp by @nastya236 in #3768
- Fix triplet_loss docstring to document the reduced output shape by @ayaangazali in #3884
- Zero-copy CPU import: mx.array(host_buffer, copy=False) on unified memory by @HaoXuAI in #3872
- Round MLX_SDPA_BLOCKS up to a multiple of 32 by @pierre427 in #3875
- Reuse Metal WAR tracking hash tables by @neilmehta24 in #3882
- Use unroll_count(4) for the NAX attention Q@K.T loop by @wyanzhao in #3843
- metal: add gemv_wide for fp16/bf16 matmuls of a few rows by @jessegross in #3888
- Fix broken docstring rendering in Linear and RNN by @ayaangazali in #3890
- Fix Adamax betas docstring and MultiOptimizer filters type by @ayaangazali in #3889
- Fix incorrect nvfp4 quantized_matmul through the split-K path by @metascroy in #3854
- [Metal] Avoid regex in custom kernel name generation by @aleroot in #3869
- Fix prod dtype promotion when reducing a size-1 axis by @eyupcanakman in #3898
- [CUDA] Fix grid overflow in gemm conv unfold kernels for >= 65,536 output positions by @AdamDLuz in #3893
- [CUDA] columnwise quantize with tma by @nastya236 in #3157
- metal: reduce NVFP4 scales per 16-lane group by @jessegross in #3934
- Update homebrew in CI by @angeloskath in #3946
- Make index autodiff errors explicitly recommend stop_gradient by @dogukanveziroglu in #3820
- Making JACCL coordinator optional by @angeloskath in #3899
- Fix docstring mismatches in the Python bindings by @ayaangazali in #3948
- Raise a clear error for an invalid quantization mode in nn layers by @ayaangazali in #3914
- Fix incorrect examples and outputs in the usage docs by @ayaangazali in #3956
- Skip test_gather_qmm_sorted cpu test on M1 mac by @zcbenz in #3973
- Fixes an axis mismatch bug in matrix norm for case -1 and 1 by @danlee2002 in #3827
- Fix BatchNorm running variance estimator by @ishtihoss in #3817
- Fix build error caused by TMA macro guard by @zcbenz in #3988
- Fix Glorot/He uniform init docstrings to label the uniform bound, not sigma by @vineethsaivs in #3831
- Fix custom metal kernel cache collision for same name, different source by @katlun-lgtm in #3833
- Fix log_cosh_loss docstring to document the element-wise loss by @winklemad in #3846
- Fix step activation docstring to match >= threshold behavior by @ayaangazali in #3902
- docs: document the reduced-precision float32 default and MLX_ENABLE_TF32 by @stoyoda0012-cyber in #3894
- Fix InstanceNorm Shape docstring to require at least 3 dimensions by @ayaangazali in #3903
- Fix filter_and_map docstring argument order for filter_fn and is_leaf_fn by @ayaangazali in #3906
- Export C++20 requirement to CMake consumers by @PhysicistJohn in #3971
- Fix implicit
threadaddress space qualifier becoming explicit in metal 4.1 by @louen in #3963 - Fix Transformer ignoring a custom encoder or decoder with no parameters by @ayaangazali in #3962
- docs: remove references to removed --no-verify-script launch flag by @latent-9 in #3959
- added eye(0) support and tests by @aaishwarymishra in #3952
- Refactor the JACCL ring and add threads for the multiple rings by @angeloskath in #3900
- Fix shapeless matmul with dynamic batch dimensions by @varshneydevansh in #3813
- Fix JIT build with old macOS SDK by @metascroy in #3853
- Fix cholesky_inv documented argument name to match the binding by @ayaangazali in #3950
- Add new_thread_unsafe_stream to the devices and streams docs by @ayaangazali in #3968
- Add gather_qqmm by @zcbenz in #3757
- Fix missing printoptions doc page from autosummary filename collision by @ayaangazali in #3985
- Document ThreadLocalStream and iinfo in the API reference by @ayaangazali in #3986
- Make "stop" optional in arange by @aaishwarymishra in #3982
- Add softsign to the nn functions docs by @ayaangazali in #3989
- Fix broken all_sum and Group references in the data parallelism example by @ayaangazali in #3996
- Fix ast.metal_kernel typo in the custom Metal kernels guide by @ayaangazali in #3997
- Fix unresolved mx.array docstring references by @ayaangazali in #3990
- python: fix bfloat16 buffer format itemsize mismatch by @reckylurker in #3975
- Template Metal complex scalar lanes by @PhysicistJohn in #3970
- Pad 2D conv input channels to reach the specialized Metal kernel by @eyupcanakman in #3904
- Support head dimension 96 in Metal full attention by @dhiltgen in #3943
- Fix mx.remainder floored-mod for float16/bfloat16 on CPU by @sashko-zakharchuk in #3976
- Fix sorted gather_mm activation row stride by @metascroy in #3960
- Fix state corruption when a primitive throws during eval by @WindChimeRan in #3675
- added complex support by @aaishwarymishra in #3984
- Validate freeze and unfreeze keys against the whole model when recursing by @ayaangazali in #3966
- Fix mlx.launch --python: flag is parsed but never forwarded to the launch script by @jonathan308 in #4002
- Fix broken fully_shard reference in FullyShardedModule docstring by @ayaangazali in #4007
- Use the current interpreter in the comparative benchmark runner by @ayaangazali in #4016
- Document MLX environment variables by @XXXXRT666 in #4000
- Fix CUDA batched GEMV grid overflow by @jasp-nerd in #3929
- Fix all_gather benchmark collapsing its input to a scalar by @ayaangazali in #4017
- Check threadgroup size in the 1-pass sdpa_vector dispatch by @apocryphx in #4018
- Align example projects with the Python 3.10 minimum by @ayaangazali in #4024
- Fix C++ benchmarks failing to build on overloaded astype by @ayaangazali in #4025
- Fix conv_transpose maxBufferLength failures on Metal via tiled unfold by @eyupcanakman in #3845
- Give each host a unique rank in Hostfile.from_list by @ayaangazali in #4027
- Fix crash when reporting partial rings in mlx.distributed_config by @ayaangazali in #4026
- docs: Do not pass MLX_METAL_FAST_SYNCH=1 by default by @katlun-lgtm in #4005
- Fix signed-integer overflow in convolution shape arithmetic by @eyupcanakman in #3938
- Template Metal C2C FFT scalar lanes by @PhysicistJohn in #3969
- Fix segfault in expand_dims for out of bounds negative axes by @Gusanidas in #4021
- Add optional dtype parameter to zeros_like and ones_like by @reckylurker in #4028
- Fix mx.longsumexp output ...
v0.32.0
What's Changed
- Generate qmm implementaions with cmake by @zcbenz in #3424
- Enable swap for all CI building CUDA by @zcbenz in #3437
- Bump minor by @angeloskath in #3438
- [CUDA] Fix qmm_naive K-tail dispatch for FP quantized kernels by @Lyxot in #3445
- Keep gguflib input-validation asserts active in release builds by @qflen in #3436
- Reuse nightly build's ccache for release by @zcbenz in #3458
- Add barrier to JACCL by @Isalia20 in #3459
- Add determinant and sign-log-determinant functions to mlx.core.linalg by @abhilashreddys in #3416
- Define ST_F8_E8M0 by @pcuenca in #3448
- Clearer error when shape dimension overflows int32 by @serenposh in #3425
- [CUDA] Fix half type matmul in cutlass kernels by @zcbenz in #3469
- Fix indexing bug in slice update with op by @angeloskath in #3483
- Make device_count() return 0 when there is no GPU by @zcbenz in #3486
- Compute contiguity from the actual occupied data by @sofinvalery in #3475
- Do not use prebuilt cpu compile preamble when headers are installed by @zcbenz in #3463
- [CUDA] Separate main loop into a function in qmm by @zcbenz in #3443
- test: Upcast the random numbers before computing their average by @sofinvalery in #3488
- Pass deployment target when linking metallib by @dhiltgen in #3501
- Fix rope single token multiple sequences by @angeloskath in #3498
- Fix qvm_split_k incorrect batch stride calculation by @angeloskath in #3497
- Fix scatter_prod GPU hang on NaN with contention by @tillahoffmann in #3492
- [jaccl] Fix race on local_staging in MeshImpl::all_reduce by @kernelpool in #3451
- Add MLX_SDPA_BLOCKS env var for 2-pass vector kernel block-count override by @adurham in #3455
- mlx launch clean by @nastya236 in #3513
- [CUDA] Fix gather_mm by @zcbenz in #3503
- [CUDA] Guard qmm_naive scale and bias loads at tile boundaries by @Lyxot in #3509
- win: fix cuda build by @dhiltgen in #3532
- Improve DLPack-compatible array imports by @XXXXRT666 in #3495
- Apply the same thread-local approach to the CPU by @angeloskath in #3537
- removed automatic prepending python in mlx launch by @nastya236 in #3536
- ci: Show stack trace on crash by @zcbenz in #3538
- Handle non-multiple-of-8 spatial dims in depthwise conv2d Metal path by @qflen in #3446
- Remove reference to groups in Conv3D extra_repr by @GriffinMB in #3559
- Add buffer caching to no_gpu CPU allocator by @dhiltgen in #3554
- Fix steel GEMM safe load offset by @sofinvalery in #3560
- Fix singleton lifetime issues at process exit by @dhiltgen in #3555
- Fix doubled-word typos in docstrings by @LeSingh1 in #3561
- Fix off_x/off_y typo in steel BaseMMAFrag::load_safe by @mdegans in #3565
- Synchronize no-GPU cache eviction with CPU streams by @dhiltgen in #3566
- Add JIT compiler support for Windows by @dhiltgen in #3556
- [Metal] Reject tensor-scale nvfp4 in qqmm by @Brooooooklyn in #3551
- Detect int32 shape-product overflow at MLX compute-shape boundaries by @qflen in #3524
- Add
copykeyword tomx.asarrayby @eyupcanakman in #3510 - Fix typos in backend code comments by @adityasingh2400 in #3581
- Implement output_shapes for GatherMM and GatherQMM by @dexwritescode in #3485
- Fix activation docstring math in log_softmax and gelu_approx by @adityasingh2400 in #3590
- docs: expose missing Python APIs by @XXXXRT666 in #3598
- Fix CUDA all-reduce planning for large inputs by @sofinvalery in #3603
- Use uv in macOS CI by @zcbenz in #3491
- Fix HDRS_LIST building when paths in header inclusion tree include spaces by @lancelotblanchard in #3607
- Enable the Metal backend by default on iOS by @eyupcanakman in #3617
- Fix int32 overflow in matvec row offset and gather-MM batch stride by @aicayzer in #3609
- Fix signed-integer overflow (UB) in roll and tile shape arithmetic by @devYRPauli in #3604
- Fix complex VJPs for log and exp by @CameronChurchwell in #3605
- added env variables to propagate width by @nastya236 in #3618
- Correct p-norm computation in triplet_loss by @pchintar in #3613
- Fix threaded compile cache cleanup by @lucasnewman in #3628
- Handle invalid dimensions in SinusoidalPositionalEncoding by @pchintar in #3615
- Correct grammatical typo in label smoothing ValueError by @madhav1k in #3641
- Roll back compile cache entry when the first trace throws by @adityasingh2400 in #3635
- Raise on arange with step == 0 instead of undefined behavior by @devYRPauli in #3640
- Fix data race in tracing state by making it thread-local by @aicayzer in #3638
- Add missing break in bool_ switch case that would result in falling through to the uint8_t case by @psolanki in #3655
- Fix axis param in nll_loss by @ishtihoss in #3651
- Fix Adafactor factored update for params with more than 2 dims by @ishtihoss in #3652
- Emit valid kernel source for non-finite float constants by @tillahoffmann in #3648
- Validate inputs in GroupNorm and InstanceNorm by @ishtihoss in #3653
- Fix flaky test_siblings_without_eval by @zcbenz in #3621
- Fix intermittent wrong bias gradient in fast::layer_norm VJP (Metal WAR hazard) by @tillahoffmann in #3630
- NAX requires setting MACOSX_DEPLOYMENT_TARGET=26.2 by @zcbenz in #3622
- Catch error in CommandBuffer and poison the events by @zcbenz in #3523
- [Improvement] CopyType::Vector in concatenate if axis=0 and contiguous by @nastya236 in #3663
- Fix Q4_1 GGUF loading by @ricky-chaoju in #3664
- Fix tree_map_with_path for namedtuples by @chrismicah in #3674
- Fix gather_qmm NAX kernel name mismatch by @scyyh11 in #3632
- Add bits and smallest_normal to finfo by @katlun-lgtm in #3679
- Fix use-after-free when a custom Metal kernel is called with different dtypes in one graph by @discobot in #3662
- Fix dimension count in orthogonal initializer error message by @Pablosinyores in #3693
- fix grid for large uncontiguous input by @nastya236 in #3666
- Ignore numpy matmul warnings in test_blas by @zcbenz in #3657
- Fix repeat zero axis shape by @chrismicah in #3698
- Cache JIT-compiled CUDA kernels by @zcbenz in #3587
- [CUDA] JIT-compile qmm_naive by @zcbenz in #3576
- Support namedtuple and tuple subclasses in tree_merge by @Pablosinyores in #3703
- Fix mx.sort vjp to transpose the permutation by @obchain in #3700
- Add isdtype, result_type, and can_cast by @katlun-lgtm in #3681
- [fix] allow exporting custom Metal kernels without a Metal backend by @xthomaswang in #3650
- Make step activation inclusive at the threshold by @Pablosinyores in #3694
- Add new_thread_unsafe_stream API by @zcbenz in #3578
- Fix JVPs of power, divmod, and slice_update with partly tra...
v0.31.2
Highlights
- Wider support for cuda quantized matmuls (#3352, #3268, #3321, #3417, #3255)
- MLX can be used by multiple threads for independent computations (#3405, #3348, #3281, #3423)
- Added CUDA FFT support
- JACCL is now a standalone lib (#3412)
What's Changed
- Bump by @angeloskath in #3244
- win: re-enable and fix cuDNN performance by @dhiltgen in #3242
- Fix crashes in multi-threaded process teardown by @louen in #3167
- [CUDA] Add FFT support by @lucasnewman in #3243
- [CUDA] Implement MaskedScatter by @Lyxot in #3151
- docs: fix PyTorch to MLX conversion example by @LxYuan0420 in #3265
- update requirements for Macbook Neo by @tosh in #3257
- fix comparison op JVP returning bool tangents instead of input dtype by @mm65x in #3253
- fix nn.GRU skipping bhn bias when hidden is None by @mm65x in #3252
- [CUDA] Pipelined QMM by @zcbenz in #3255
- tests: harden memory leak check in test_siblings_without_eval by @booxter in #3088
- Slice update with operation by @angeloskath in #3266
- Nax Refactor by @jagrit06 in #3271
- Fix building with CUDA toolkit 13.2 by @zcbenz in #3273
- [CUDA] fp and int4 quants for qmm_sm80 by @zcbenz in #3268
- Fix repr of conv layers by @angeloskath in #3275
- Merge DeviceStream into CommandEncoder by @zcbenz in #3264
- [CUDA] Search system-installed CUDA toolkit for headers by @zcbenz in #3277
- Create default random key lazily by @zcbenz in #3278
- Support indexing with any type which implmented
__index__by @aisk in #3210 - Fix sort NaN handling for float16 and bfloat16 by @Lyxot in #3269
- Use thread local storage for frontend compile cache by @zcbenz in #3280
- [Metal][Performance]: Add split-K for quantized matmul (small M) by @Ziqiao-git in #3120
- [Metal] Fix depthwise conv 1D kernel name for large variant by @Brooooooklyn in #3289
- Fix stale transform copy-chain leaks by @Brooooooklyn in #3290
- Implement Pad::vmap to replace NYI stub by @Aristide021 in #3304
- logo files by @andresy in #3308
- Fix vmap + floor_divide: preserve integer dtype by @robert-johansson in #3292
- Fix moved-from shape bug in broadcast_arrays causing vmap bus error by @Aristide021 in #3310
- Use nb::ndarray for checking arrays by @zcbenz in #3283
- Add output_shapes for AddMM by @pHequals7 in #3262
- Manage Metal objects with smart pointers by @zcbenz in #3282
- [CUDA] support sorting complex numbers by @Lyxot in #3286
- Add norm parameter to FFT transforms (backward/ortho/forward) by @Aristide021 in #3287
- Make each thread have its own default stream by @zcbenz in #3281
- [CUDA] Implement BlockMaskedMM by @Lyxot in #3299
- Fix np bfloat16 misinterpreted as complex by @kellen-sun in #3146
- Remove no longer needed const_cast by @zcbenz in #3325
- Bump actions/deploy-pages from 4 to 5 by @dependabot[bot] in #3334
- Fix use after move by @angeloskath in #3343
- Decouple CommandEncoder from Device by @zcbenz in #3316
- Add vmap for BroadcastAxes by @angeloskath in #3344
- Add fftfreq, rfftfreq and scalar axes for fftshift/ifftshift by @declanhealy2 in #3298
- [Metal] Support sorting complex numbers by @Lyxot in #3314
- [CUDA] Fallback QMM by @zcbenz in #3315
- Make CommandEncoder thread local by @zcbenz in #3348
- [CUDA] 3/5/6-bit quants for qmm_naive by @zcbenz in #3352
- Fix regression in array creation by @angeloskath in #3353
- Use
metalas the front-end for the metal linker by @louen in #3354 - Add printoptions by @ChristophePRAT in #3333
- Add a convenience for making local streams in python by @angeloskath in #3355
- Fix CMake finding wrong Python during pip install by @fijimunkii in #3375
- [CUDA] Add GatherQMM for quantized gather matmul by @Lyxot in #3321
- fix: fail build when Metal compiler header resolution fails by @dogukanveziroglu in #3332
- Fix: Correct cross-attention query routing in Post-LN TransformerDecoderLayer by @suryawanshishantanu6 in #3382
- [CUDA] Thread safety by @zcbenz in #3367
- Fix test "test get streams" missing initialization by @dseredkin in #3376
- Conjugate VJP and JVP support by @CameronChurchwell in #3386
- Fix int16 overflow in SDPA NAX mask indexing for KV sequences > 32K by @Clydingus in #3361
- Avoid joining threads on exit by @zcbenz in #3388
- Add clear_streams API for cleanup before exit by @zcbenz in #3395
- Update nanobind version to v2.12.0 by @jrp2014 in #3396
- Jaccl refactor by @angeloskath in #3412
- Fixes for CUDA CI by @zcbenz in #3413
- Validate safetensors data offsets by @MillaFleurs in #3364
- Validate safetensors data offsets against file boundaries by @matinsaurralde in #3410
- Document sort stability and NaN handling by @NeuralNoble in #3400
- ThreadLocalStream in C++ by @zcbenz in #3405
- Fix jaccl init bug by @angeloskath in #3418
- Segmented mm nax kernel by @angeloskath in #3419
- [CUDA] gather_mm by @zcbenz in #3414
- [CUDA] GatherQMM matrix-matrix sm80/naive path by @Lyxot in #3417
- [CUDA] Handle residue k in qmm_naive by @zcbenz in #3379
- Speed up NAX split-K by better tuning and routing and fix NAX addmm by @angeloskath in #3422
- Make Scheduler::enqueue thread safe by @zcbenz in #3423
- Fix flaky TestVmap.test_vmap_masked_scatter by @zcbenz in #3421
- Fix synchronize for ThreadLocalStream by @angeloskath in #3429
- Fix bytes_per_key truncation in random kernels (Metal + CUDA) by @dogukanveziroglu in #3432
- Throw meaningful error when Metal device is not found by @dogukanveziroglu in #3428
- Fix kernel cache collision in Compiled constructor by @dogukanveziroglu in #3427
- Fix mx.prod vjp for complex types by @CameronChurchwell in #3433
New Contributors
- @LxYuan0420 made their first contribution in #3265
- @tosh made their first contribution in #3257
- @mm65x made their first contribution in #3253
- @booxter made their first contribution in #3088
- @Ziqiao-git made their first contribution in #3120
- @Brooooooklyn made their first contribution in #3289
- @Aristide021 made their first contribution in #3304
- @pHequals7 made their first contribution in #3262
- @declanhealy2 made their first contribution in #3298
- @fijimunkii made their first contribution in #3375
- @dogukanveziroglu made their first contribution in #3332
- @suryawanshishantanu6 made their first contribution in #3382
- @dseredkin made their first contribution in #3376
- @CameronChurchwell made their first contribution in https://gith...
v0.31.1
What's Changed
- Bump the patch version by @angeloskath in #3185
- Skip Hopper-only kernels in CI by @zcbenz in #3184
- [CUDA] Fsdp (easy) by @nastya236 in #3130
- Fix ref leak in mx.save/load with file like object by @aisk in #3187
- Fix/missing libs in docs by @ChristophePRAT in #3190
- feat: adding the bartlett function by @Vlor999 in #3155
- Bump actions/download-artifact from 7 to 8 by @dependabot[bot] in #3189
- Bump actions/upload-artifact from 6 to 7 by @dependabot[bot] in #3188
- [CUDA] Quantized GEMV by @zcbenz in #3180
- [CUDA] Use fp16 accumulation for 4-bit quant in GEMV by @zcbenz in #3197
- [CUDA] implement Hadamard transform by @Lyxot in #3179
- Improve mlx.distributed_config by @angeloskath in #3199
- PR#3226 Fix by @MillaFleurs in #3227
- PR #3220 LayerNorm VJP returns zeros_like(weight) instead of zeros_like(bias placeholder) by @MillaFleurs in #3231
- [CUDA] Faster compilation and batch support in QMV by @zcbenz in #3213
- Validate num_splits in split by @MillaFleurs in #3234
- Fix return value in einsum_path for simple contractions by @MillaFleurs in #3232
- Validate dims in rope by @MillaFleurs in #3230
- Fix assigning bool to float16/bfloat16 by @MillaFleurs in #3229
- Fix load_weights with strict=False to filter extra weights before update by @gmin7 in #3214
- Remove custom fp4/fp8 classes by @zcbenz in #3212
- [CUDA] Support 3/5/6-bit quants in QMV by @zcbenz in #3236
- Hybrid sharding by @nastya236 in #3194
- win: fix cuda build by @dhiltgen in #3204
- Remove quantized_utils.cuh by @zcbenz in #3237
- [CUDA] Implement SegmentedMM by @Lyxot in #3238
- Add initial tuning for M5 pro and max by @jagrit06 in #3211
- [CUDA] Use qmv kernel for fp quantizations by @zcbenz in #3239
New Contributors
- @ChristophePRAT made their first contribution in #3190
- @Lyxot made their first contribution in #3179
- @gmin7 made their first contribution in #3214
Full Changelog: v0.31.0...v0.31.1
v0.31.0
Highlights
- Initial version of QMMs for CUDA (#3160)
- JACCL mesh bandwidth improvements (#3174)
- Massive speedups for 3D convs (#3147)
- Continued improvements to qqmm (#3106, #3022)
What's Changed
- Patch bump by @angeloskath in #3102
- is_available() should check the device index too by @andresy in #3107
- Fix residency set with user provided buffer by @awni in #3108
- Cleanup test_fast_sdpa.py by @zcbenz in #3112
- [CUDA] Set current device before allocating memory by @zcbenz in #3110
- Quantize module to QQLinear by @nastya236 in #3106
- [CUDA] Use cuDNN SDPA for decoding when using fixed-size KV cache by @zcbenz in #3113
- register pressure by @nastya236 in #3116
- Fix precision in Metal fused attention by @awni in #3119
- [CUDA] Attention sinks in cuDNN SDPA by @zcbenz in #3118
- Fix donation in sdpa vector by @angeloskath in #3121
- Manage stream placement in import function by @awni in #3127
- fix: propagate quantization mode in QuantizedAllToShardedLinear / QuantizedShardedToAllLinear by @vskiwi in #3133
- [featuring] - add hanning window function by @Vlor999 in #3124
- feat: adding the hamming function by @Vlor999 in #3135
- Tensor scale nvfp4 by @nastya236 in #3022
- Fix fence synchronization accross command buffers by @awni in #3144
- Export: preserve Dtype state values in export callback arguments by @skryl in #3145
- [Metal] Fix 32-bit integer overflow in conv3d unfold kernel by @kellen-sun in #3143
- [Metal][Performance] Add implicit matmul pathway for mx.conv3d by @belkakari in #3147
- [Metal] Fix event leak by @awni in #3159
- [CUDA] FPxINT quantized matmul for Hopper by @zcbenz in #3160
- feat: implement mlx.core.blackman by @Vlor999 in #3136
- Enable setting thread block cluster for Hopper and later by @zcbenz in #3168
- [CUDA][NCCL] group split by @nastya236 in #3172
- JACCL refactor and small update by @angeloskath in #3174
- [CUDA] Heuristics for Hopper QMM by @zcbenz in #3173
- Fix compile_fuse broadcast split aliasing bug by @robert-johansson in #3166
- Enable passing in a GPU architecture string via env var by @angeloskath in #3176
- Bump the minor version by @angeloskath in #3183
New Contributors
- @vskiwi made their first contribution in #3133
- @Vlor999 made their first contribution in #3124
- @skryl made their first contribution in #3145
- @kellen-sun made their first contribution in #3143
- @belkakari made their first contribution in #3147
- @robert-johansson made their first contribution in #3166
Full Changelog: v0.30.6...v0.31.0
v0.30.6
Highlights
- Much faster bandwidth with JACCL on macOS >= 26.3 (some numbers)
What's Changed
- patch by @awni in #3093
- Disable managed memory on WSL when concurrentManagedAccess is not supported by @jessegross in #3095
- Fix non simd f16 build by @awni in #3097
- Fix 2pass sdpa on < M2 by @awni in #3099
- JACCL update by @angeloskath in #3094
- Fix qmv_impl for small N by @manuelcandales in #3096
- Patch for multi device CUDA by @awni in #3100
New Contributors
- @manuelcandales made their first contribution in #3096
Full Changelog: v0.30.5...v0.30.6
v0.30.5
What's Changed
- patch by @awni in #3074
- [CUDA] Fallback Event impl when there is no hardware cpu/gpu coherency by @zcbenz in #3070
- Tune CUDA gaph sizes on B200 and H100 by @awni in #3077
- [Docs] Simple example of using MLX distributed by @stefpi in #2973
- Use lower-right causal mask alignment consistently by @Anri-Lombard in #2967
- Fix ALiBi slopes for non-power-of-2 num_heads by @vovw in #3071
- More useful error for large indices by @awni in #3079
- Fix nax condition for iphone by @awni in #3083
- Fallback to pinned host memory when managed memory is not supported by @zcbenz in #3075
- Fix failing python tests on Windows by @zcbenz in #3076
- [Metal] Tune splitk gemm dispatch conditions and partition sizes by @awni in #3087
- Fix for NAX overflow. by @awni in #3092
New Contributors
Full Changelog: v0.30.4...v0.30.5
v0.30.4
Highlights
- Metal: Much faster vector fused grouped-query attention for long context
- CUDA: Several improvements to speed up LLM inference for CUDA backend
- CUDA: Support for dense MoEs
- CUDA: Better support for consumer GPUs (4090, 5090, RTX 6000, ...)
What's Changed
- patch bump for next release by @awni in #2991
- Fix fence by @awni in #2998
- Reverts changing the MLX_IBV_DEVICES to MLX_JACCL_DEVICES by @angeloskath in #2999
- fix distributed all_to_sharded bias shard axis from -2 to -1 by @gufengc in #2987
- Fix sharding of quantized models with non-power-of-2 bits by @kernelpool in #3006
- Update CCCL to v3.1.3 by @zcbenz in #3012
- Fix python package install path in stubgen by @zcbenz in #3009
- Type Enhancement for Func Transforms and Bug Fix by @XXXXRT666 in #3003
- Do not clear disk space in setup-linux by @zcbenz in #3013
- Do not give workflow boolean inputs default values by @zcbenz in #3014
- Fix negative dim indexing by @MillaFleurs in #2994
- Windows CI by @zcbenz in #3021
- Optimize erf function with expm1f in Metal backend by @bjornefisk in #3025
- [CUDA] Faster grouped mm by @zcbenz in #3011
- PR 3007 Fix Seg Fault by @MillaFleurs in #3008
- Use higher precision for linspace with double by @awni in #3029
- Handle data smaller than BUFFER_SIZE in jaccl recv by @rltakashige in #3033
- build 26.0 release in actions by @awni in #3035
- Remove xmlrunner from macOS CI by @zcbenz in #3032
- Columnwise quantize by @nastya236 in #2989
- Turn nccl_stub into a normal target by @zcbenz in #3037
- Use cuda::std for math ops by @zcbenz in #3041
- win: symbol exports and minor fixes by @dhiltgen in #3024
- CUDA gather mv by @angeloskath in #3039
- Link with prebuilt OpenBLAS and fix shared libs build on Windows by @zcbenz in #3036
- Allow take on empty array when it makes sense by @awni in #3046
- Add missing include to buffer_cache.h by @Anri-Lombard in #3053
- Build and test python package on Windows CI by @zcbenz in #3049
- Fix some MSVC compilation errors by @zcbenz in #3048
- Use C++20 by @zcbenz in #3050
- Faster two pass sdpa by @awni in #3023
- Find system-installed cuDNN on Windows by @zcbenz in #3052
- Fix some NVCC warnings when building CUDA backend with MSVC by @zcbenz in #3038
- Hide symbols by default for mac/linux by @zcbenz in #3057
- [CUDA] Fast sorting by @awni in #3060
- Fix flaky macOS test by @awni in #3063
- Update pre-commit hooks and versions for clang-format, black, and isort by @NripeshN in #3059
- GPU discovery by @dhiltgen in #3055
- Add NAX Split-K GEMM for large-K matmuls to improve performance by @hxu296 in #3018
- Improve CPU discovery by @dhiltgen in #3068
- Fix long cache file path on Windows by @zcbenz in #3065
- Better support consumer CUDA GPUs by @jessegross in #3056
- Delay load CUDA libs and resolve DLL paths at runtime by @zcbenz in #3061
- Do not require ConcurrentManagedAccess when not used by @zcbenz in #3062
- Fp qmv by @awni in #2984
- remove thrust by @awni in #3067
New Contributors
- @gufengc made their first contribution in #2987
- @kernelpool made their first contribution in #3006
- @bjornefisk made their first contribution in #3025
- @rltakashige made their first contribution in #3033
- @dhiltgen made their first contribution in #3024
- @hxu296 made their first contribution in #3018
- @jessegross made their first contribution in #3056
Full Changelog: v0.30.3...v0.30.4