Text-to-Audio
Transformers
Diffusers
Safetensors
ACE-Step
feature-extraction
audio
music
text2music
custom_code
Instructions to use ACE-Step/Ace-Step1.5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ACE-Step/Ace-Step1.5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-audio", model="ACE-Step/Ace-Step1.5", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ACE-Step/Ace-Step1.5", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Field report: ACE-Step 1.5 XL-Turbo on an 8 GB AMD card β the "lead rule", melody guidance and a blind test vs Stable Audio 3
#31
by charlysporty - opened
I've been running ACE-Step 1.5 locally for a few weeks (acestep.cpp, Vulkan, Radeon RX 6600 8 GB, Windows 11) and wrote up what I learned, in German and English:
https://huggingface.co/spaces/charlysporty/lokale-musikerzeugung
A few findings that might be useful to others:
- Every melody needs a lead. Instrumentals without a named lead scored 100 % on my scale-adherence metric but sounded "not really harmonious". With a lead instrument they were fine β but only if the lead instrument comes first in the caption. Otherwise the first instrument named in the template takes over (e.g. the harmonica in a Chicago blues caption beat the guitar I had chosen).
- Melody guidance via the cover task:
audio_cover_strength0.3 with a deliberately plain, dry synth source gave the best tone and phrasing; from 0.5 the synth timbre bleeds through. - Blind test vs Stable Audio 3 (both instrumental, same lead, seed, tempo and loudness, 5 styles): the listener preferred ACE-Step 4 times, once a tie, Stable Audio never β although all ten versions were rated "very good".
- Performance: XL-Turbo renders 120 s in about 210 s (up to 6.5 GB VRAM), an 8-minute song in about 17 minutes.
- What did not work: register words ("low register"), mix words ("subtle bass"), negations, and the
completetask with the base model (phantom vocals).
Feedback, counter-examples and your own measurements are very welcome β especially whether the lead rule holds for other listeners and hardware.
(Programming, measurements and write-up were done with Claude as an AI assistant; listening judgements are mine.)