YuE2 Local Guide: Setup, Hardware and Lyrics-to-Song Results
2026/10/05

YuE2 Local Guide: Setup, Hardware and Lyrics-to-Song Results

Run YuE2 locally: official downloads, Linux GPU setup, Mac Q8 benchmark results, lyrics workflow and an evidence-based comparison with ACE-Step.

YuE2 turns lyrics and a style description into a song with vocals and accompaniment. For local use, distinguish the official unquantized Python route from a GGUF / audio.cpp route. Their hardware requirements, files and available controls differ. YuE2-3B's current model card lists CC-BY-NC-4.0, so review the non-commercial restriction before choosing it for a project.

Checked October 5, 2026. The Mac results below come from our September 23 benchmark, not a new test of every installer or computer. For a wider overview, see open-source music models.

Downloads and hardware: choose the correct route

RouteFiles / backendHardware evidence
Official YuE2-3BPython pipeline and official model assetsModel card: Linux, Python 3.10+, 24 GB NVIDIA GPU with BF16 support
Local AI Music catalogYuE2-3B Q8 GGUF, VAE and configuration sidecars; audio.cppHistorical test: M5 Pro, 64 GB unified memory, Metal; about 7.0 GiB peak footprint
A different quantization or Windows buildDepends on the backend and releaseNot measured here; check that exact package

Get official files from m-a-p/YuE2-3B. The desktop catalog uses audio-cpp/Yue2-3B-GGUF, a separate converted distribution. Do not mix its GGUF files into the official Python loader. The app's listed download size—about 4.6 GB—is not a VRAM requirement. The 24 GB official GPU recommendation is not a claim that our quantized Mac route needs 24 GB VRAM.

Official local setup

In an isolated Python environment on a supported Linux/NVIDIA machine, the model card currently gives this package installation:

python -m pip install huggingface-hub==0.36.2
hf download m-a-p/YuE2-3B yue2_infer-0.1.5-py3-none-any.whl --local-dir .
python -m pip install ./yue2_infer-0.1.5-py3-none-any.whl

The official pipeline can save both the song and its generation artifacts. This short example uses new example lyrics, not the benchmark input and not a claimed finished output:

from yue2 import YuE2Pipeline

pipe = YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda")
song = pipe(
    style="Acoustic folk, warm vocal, fingerpicked guitar, gentle drums",
    lyrics="[verse]\nA window catches morning light\nThe road unfolds beyond the night\n[chorus]\nWe carry little sparks of day\nAnd sing them softly on our way",
    cot="full",
    seed=11,
)
song.save("first-song.flac")
song.save_artifacts("outputs/first-song")

The first run fetches model assets. Keep the resulting settings and score with the audio so that later edits have a traceable starting point. Check the current model card if package filenames or APIs change. We have verified this example against its documented interface, not executed the official CUDA pipeline in this update.

Mac desktop workflow

The current Local AI Music model catalog includes YuE2 Q8 through audio.cpp. In your installed version, confirm that YuE2 appears in the music model list, read the license, download the complete package, and start with short lyrics. Choose one language and a coherent style description; save the input with the generated file.

Use a verse and chorus before attempting a full song. If a line is skipped, shorten or restructure it and compare runs with the same seed where the interface exposes one. Our tested YuE2 route follows the lyrics rather than a precise duration setting. Do not expect “two minutes” in a prompt to guarantee a two-minute output.

Official score-editing capabilities are not a promise that every desktop UI exposes the same controls. Check the release you install. For lyric structure, see turning lyrics into songs.

Mac results: YuE2 versus ACE-Step

Download the historical measurement summary CSV. This is an extract of the September 23 record, not a new benchmark.

September 23, 2026: MacBook Pro M5 Pro, 64 GB unified memory, audio.cpp 487800f5, Metal. Each music run started in a separate process; elapsed time includes model loading. The tests used original project lyrics and seed 11.

Model and caseOutput lengthElapsed timePeak memory footprint
YuE2 3B Q8, Chinese city pop144 seconds166 seconds7.0 GiB
YuE2 3B Q8, English rock120 seconds130 seconds6.9 GiB
YuE2 3B Q4, Chinese city pop132 seconds109 seconds5.6 GiB
ACE-Step 1.5 Turbo Q8, Chinese city pop120 seconds89 seconds7.9 GiB
ACE-Step 1.5 Turbo Q8, English rock90 seconds73 seconds7.7 GiB

The Chinese test used a female-vocal city-pop description with electric piano, clean funk guitar and a 104 BPM target. The opening verse was:

[verse]
晚风轻轻掠过长街
霓虹倒映熟悉的脸
我把沉默写成明信片
等一句回答穿过时间

The full test also included a pre-chorus, repeated chorus and bridge. This excerpt alone will not reproduce the 144-second result. The requested style and tempo are input intentions, not measurements of the finished track.

We transcribed generated songs with Qwen3-ASR-1.7B: Chinese character error was 3.8% for YuE2 Q8, 13.0% for Q4, and 7.7% for ACE-Step. ASR itself makes mistakes on singing; this is a small comparison of lyric adherence, not a music-quality score. YuE2's English output added an unrequested opening line in this sample. There was no YuE2 instrumental test.

Listen before choosing

These are existing YuE2 product showcase tracks, not the timed benchmark outputs above. They help assess vocals and arrangement, while the table explains one measured runtime configuration.

“晚风吹过旧车站”:

“门前的老槐树”:

Which model should you choose?

  • Try YuE2 Q8 for non-commercial lyric-focused experiments. It preserved more Chinese lyrics in this small test. Check the current model license and any separate permissions needed for your use.
  • Start with ACE-Step for a faster drafting loop or duration control. The official ACE-Step project uses MIT licensing; the exact model and any reference material still need review. Our sample does not prove it will always be faster.
  • Do not select Q4 solely by download size. It reduced memory here but lost lyric accuracy. Listen to the result before accepting the trade-off.

Troubleshooting and next steps

Missing VAE or configuration files can break a converted model package: download the complete backend-specific set. For memory failures, close other models and shorten the task before changing quantization. Keep a copy of the unedited export; check clipping and loudness in an audio editor before distribution. We observed peaks reaching 0 dBFS in the historical tests, which merits checking rather than an automatic “mastered” label.

A desktop installer and a model license are separate things. Paying for an installer does not remove YuE2's non-commercial condition, and local generation does not settle rights in lyrics or reference recordings. For a local-versus-cloud decision, read the offline Suno alternative; for arranging your input, see writing songs with AI offline.

Newsletter

Join the community

Subscribe to our newsletter for the latest news and updates