Choose theme
AI Resources
ACE-Step 1.5
ACE-Step 1.5 is a locally runnable music model for generating songs and revising audio, with reference-guided creation, repainting and model-dependent stem tools.
A language model can plan the song and a diffusion model produces audio. The project offers 2B and XL 4B diffusion variants, a local interface and API, plus hosted demonstration paths. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Draft a song, then revise the audio
A creator with a rough track can regenerate a selected passage or guide a new result with reference audio. This gives a revision path beyond replacing the whole song with another text prompt.
Why it stands out
Planning and audio generation are separate
The language model supplies song planning and rewriting; the diffusion model produces audio. The documented low-memory DiT-only path disables the language model, so the smallest setup does not include the same planning stage.
Availability
Local models with different requirements
The repository documents Python 3.11–3.12, local UI/API paths and several hardware backends. The 2B baseline and XL 4B variants have different memory requirements; a hosted demo's access and data handling are separate from running locally.
Why it matters
What makes it useful
If one passage needs changing while the rest of a draft should remain the starting point, repainting targets a selected region. Listen across its boundaries afterward: the project's limitations include rough transitions and inconsistent results across seeds and durations.
What to know
Where it fits
It generates and reshapes source audio before a wider arranging or mixing workflow. Covers, accompaniment, style adapters and stem tasks depend on the selected model and feature; this is not a full digital audio workstation.
Notable points
What stands out
The v0.1.8 release adds Retake variation generation. Its retake_variance control mixes fresh noise into the diffusion start on a 0–1 scale, providing an explicit variation control rather than changing the prompt alone.
Before using
What to review
Choose the model by both task support and memory. The project's base, SFT and turbo variants do not expose every task in the same way.
The 2B guide lists at least 4 GB VRAM for DiT-only and 6 GB for language-model-plus-DiT use. XL needs more memory; CPU inference is supported but described as substantially slower.
The project reports inconsistent generations, coarse vocals and limited fine control. Its performance and quality claims are not an independent LifeHubber listening test.
The project's current terms are linked on its repository.
Reader fit
Who may find it relevant
Musicians experimenting with generated drafts, reference audio and selective revision on their own hardware.
Builders connecting a local music interface or API to a broader production workflow.
Someone wanting a finished hosted service or complete arranging and mixing tools may need a different setup.
Editorial note
Why LifeHubber lists it
Stem extraction and adding a musical layer need the right checkpoint. In the project's task matrix, Extract, Lego and Complete are supported by the base variants, while the SFT and turbo variants do not support those tasks. Pick that capability before choosing a model for its generation speed.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
More in Music / Image Gen Models
Keep browsing this category
Explore more media-generation model resources.
CoMoVi
IGL-HKUST/CoMoVi
A framework for co-generating 3D human motion and realistic videos, with a focus on motion-conditioned video generation and training workflows.
Qwen-Image-2.1
QwenLM/Qwen-Image-2.1
Qwen's open-weight image model for text-to-image generation, single- and multi-reference editing, marked-region changes, subject extraction, native transparent RGBA output, and 2K workflows.
PersonaLive
GVCLab/PersonaLive
A portrait image-animation framework for live-streaming-style video generation research, with offline and online inference, pretrained weights, a Web UI, and acceleration notes.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.