mlx-serve/Deep dive · Music generation

Type a vibe.
Get a song.

"Upbeat synthwave with driving bass" — and ACE-Step 1.5 XL Turbo composes an original 48 kHz stereo track in 8 diffusion steps, entirely on your Mac. Paste lyrics and it sings them; leave them out for an instrumental. No cloud, no subscription, no usage caps. And when the voice is the point, MiniMax Music 3 sings your lyrics for up to six minutes.

Download MLX Core How it works
8 diffusion steps per track 10 s – 10 min per generation 48 kHz stereo output
How it works

Describe. Generate. Keep every take.

Open the Audio pane's Music tab, describe the style — genre, mood, instrumentation — and generate. Want vocals? Paste lyrics (verse/chorus tags steer the arrangement) and pick the vocal language. Want control? Pin the BPM, key, and time signature; leave them free and the model chooses.

Style-prompt starters and original lyric templates ship in the Examples menus, and you can save your own to reuse. Every track lands in a persistent history — play, stop, reveal in Finder — with a settings file beside it recording the exact prompt, lyrics, and parameters, so any take is reproducible.

MLX Core Audio pane, Music tab — style prompt, lyrics, duration and musical steering controls with a generation history
The Music tab: style prompt, optional lyrics, musical steering — and a history of every take.
Under the hood

A music diffusion model, ported natively

ACE-Step 1.5 XL Turbo runs on the same no-Python engine as everything else here: a text encoder conditions a 32-layer diffusion transformer, an audio VAE decodes to waveform, and the turbo schedule needs just 8 steps end to end. The port is validated against the reference implementation with per-component numerical oracles.

One click downloads the converted model (~6.3 GB); it generates comfortably in about 9 GB of memory, loads on demand, and frees itself afterwards — chat and music coexist in one server process.

  • No subscription, no credits — generate all night on your own silicon.
  • Your lyrics stay yours — nothing you write leaves the machine.
  • Works offline — once the model is downloaded, airplane mode is fine.
MiniMax Music 3

When the voice is the point

MiniMax Music 3 writes full songs: an 8B language model composes the track frame by frame from your style caption and lyrics, then a diffusion decoder renders it at 44.1 kHz. It is the strongest vocal model in the lineup, with songs up to six minutes. Lyrics are required, and structure tags like [verse] and [chorus] go on their own lines; tempo, key and meter live in the caption rather than in controls.

Ask the chat for a song and the bundled music3 skill writes the three-block caption format the model was trained on, plus original lyrics, instead of a one-liner. ACE-Step stays the fast option at 8 steps.

  • Sings your lyrics — the strongest vocals shipped here.
  • Songs up to 6 minutes at 44.1 kHz.
  • 13.6 GB download, about 20 GB of memory to run.
  • music3 chat skill writes the trained caption format for you.

Music that plugs into everything else

API

One endpoint for developers

POST /v1/audio/music-generations with a prompt, optional lyrics, duration, and musical steering — get a stereo WAV back, with per-step SSE progress. MiniMax Music 3 uses the same endpoint; there lyrics is required. Full API docs.

Video

Score your generated videos

Feed a generated track into audio-to-video and the visuals perform to your soundtrack — beat, mood, and all.

Voice

Sibling tab: cloning & TTS

The Audio pane's Voice tab does zero-shot voice cloning — six seconds of audio, no transcript, bit-for-bit validated.

Reproducible

Every take documented

A .txt sidecar beside each track records the model, prompt, lyrics, and parameters — recreate or iterate on any result later.

Local music generation, answered

What hardware do I need?

An Apple Silicon Mac with 16 GB works well — the 8-bit model generates in about 9 GB of memory and is a ~6.3 GB one-click download.

Can it sing my lyrics?

Yes. Paste lyrics and the vocal performance follows them — structure tags like verse and chorus steer the arrangement, and the vocal language is selectable. Empty lyrics produce an instrumental.

How much control do I get over the music itself?

Style, mood, and instrumentation come from the prompt; BPM (30–300), key/scale, and time signature can be pinned exactly or left to the model. Track length runs from 10 seconds to 10 minutes, and a seed makes any generation repeatable.

ACE-Step or MiniMax Music 3: which do I pick?

ACE-Step is the fast one: 8 diffusion steps, instrumental or vocal, BPM/key/meter pinnable, ~6.3 GB. MiniMax Music 3 is the vocal one: it composes frame by frame and sings your lyrics for up to six minutes, at 13.6 GB down and about 20 GB of memory. Lyrics are required on Music 3, and the tempo/key controls don't exist there; put those facts in the caption instead.

Who owns the output?

It's generated on your machine from your prompt — there's no service holding rights over your renders. ACE-Step is an open-weights model; check its license for the fine print on commercial use of outputs.

More deep dives

Your studio. Your silicon. No meter running.

Download MLX Core, grab the music model with one click, and render as many takes as you like.