Skip to content

Releases: ml-explore/mlx

v0.32.3

Choose a tag to compare

@zcbenz zcbenz released this 29 Sep 00:37
64ea011

What's Changed

New Contributors

Read more

v0.32.2

Choose a tag to compare

@zcbenz zcbenz released this 25 Aug 10:58
1f8e74e

What's Changed

New Contributors

Full Changelog: v0.32.1...v0.32.2

v0.32.1

Choose a tag to compare

@zcbenz zcbenz released this 18 Aug 03:45
3a62199

What's Changed

Read more

v0.32.0

Choose a tag to compare

@angeloskath angeloskath released this 07 Jul 22:29
7a1d4f5

What's Changed

Read more

v0.31.2

Choose a tag to compare

@angeloskath angeloskath released this 22 Apr 01:40
68cf2fd

Highlights

What's Changed

New Contributors

Read more

v0.31.1

Choose a tag to compare

@angeloskath angeloskath released this 12 Mar 06:58
ce45c52

What's Changed

New Contributors

Full Changelog: v0.31.0...v0.31.1

v0.31.0

Choose a tag to compare

@angeloskath angeloskath released this 28 Feb 06:22
365d6f2

Highlights

  • Initial version of QMMs for CUDA (#3160)
  • JACCL mesh bandwidth improvements (#3174)
  • Massive speedups for 3D convs (#3147)
  • Continued improvements to qqmm (#3106, #3022)

What's Changed

New Contributors

Full Changelog: v0.30.6...v0.31.0

v0.30.6

Choose a tag to compare

@angeloskath angeloskath released this 06 Feb 17:05
185b06d

Highlights

  • Much faster bandwidth with JACCL on macOS >= 26.3 (some numbers)

What's Changed

New Contributors

Full Changelog: v0.30.5...v0.30.6

v0.30.5

Choose a tag to compare

@awni awni released this 03 Feb 02:56
adcbb91

What's Changed

  • patch by @awni in #3074
  • [CUDA] Fallback Event impl when there is no hardware cpu/gpu coherency by @zcbenz in #3070
  • Tune CUDA gaph sizes on B200 and H100 by @awni in #3077
  • [Docs] Simple example of using MLX distributed by @stefpi in #2973
  • Use lower-right causal mask alignment consistently by @Anri-Lombard in #2967
  • Fix ALiBi slopes for non-power-of-2 num_heads by @vovw in #3071
  • More useful error for large indices by @awni in #3079
  • Fix nax condition for iphone by @awni in #3083
  • Fallback to pinned host memory when managed memory is not supported by @zcbenz in #3075
  • Fix failing python tests on Windows by @zcbenz in #3076
  • [Metal] Tune splitk gemm dispatch conditions and partition sizes by @awni in #3087
  • Fix for NAX overflow. by @awni in #3092

New Contributors

Full Changelog: v0.30.4...v0.30.5

v0.30.4

Choose a tag to compare

@zcbenz zcbenz released this 27 Jan 22:27
2f324cc

Highlights

  • Metal: Much faster vector fused grouped-query attention for long context
  • CUDA: Several improvements to speed up LLM inference for CUDA backend
  • CUDA: Support for dense MoEs
  • CUDA: Better support for consumer GPUs (4090, 5090, RTX 6000, ...)

What's Changed

New Contributors

Full Changelog: v0.30.3...v0.30.4