SoulX-Singer

B
B tier on AI Song Cover GeneratorsScore 7.0 · #6 of 55
Android app
Not listed
Free plan
Yes
Runs on
Linux, self-hosted, Web
github.com
The SoulX-Singer homepage

Summary

SoulX-Singer is an open-source model for generating singing voices for singers it has not been tuned on. It can control pitch, rhythm, and expression using either melody-conditioned F0 contours or score-conditioned MIDI notes. Its voice-conversion system changes raw singing audio to a target singer’s voice while retaining melody, rhythm, and lyrics, without requiring lyric or MIDI transcriptions. The system supports Mandarin Chinese, English, and Cantonese, as well as singer-timbre cloning, cross-lingual synthesis, and lyric editing while retaining natural prosody. Preprocessing tools handle vocal separation, dereverberation, F0 extraction, voice activity detection, and lyric and note transcription. A MIDI editor can be used to adjust lyrics, phoneme alignment, pitches, and durations before synthesis. Soul-AILab provides singing-generation and vocal-conversion demos, plus a separate MIDI editor, through Hugging Face Spaces. The repository also supports local inference with Conda, Python 3.10, pip-installed dependencies, and web UI scripts. Code and weights use the Apache 2.0 license. Maintainers caution that automatic preprocessing may misalign singing with lyrics and notes, so manual correction may be needed.

Who it is for

SoulX-Singer is suited to researchers and developers working with singing synthesis, voice conversion, or MIDI-based control. It offers both online demos and a local deployment route, but automatic preprocessing may need correction.

What is good

  • Generates voices for unseen singers without fine-tuning.
  • Controls pitch, rhythm, and expression with F0 or MIDI.
  • Supports Mandarin Chinese, English, and Cantonese.
  • Provides vocal separation and singing transcription tools.
  • Code and model weights use the Apache 2.0 license.

What to know first

  • Automatic preprocessing may misalign audio, lyrics, and notes.
  • Local inference uses Conda, Python 3.10, and pip dependencies.
  • Unauthorized impersonation and deceptive audio are prohibited.

Verdict

SoulX-Singer offers singing synthesis and conversion with several control and editing options, along with demos and local inference. Plan to review preprocessing results and respect the project’s consent and impersonation restrictions.

SoulX-Singer plans and pricing

All plans
SoulX-Singer Free Researchers and developers are free to use the code and model weights · Apache 2.0 license github.com · 1 Oct 2026

Compared on AI song cover generators

Voice cloning
Yesgithub.com
Vocal input
Yesgithub.com
MIDI support
Yesgithub.com
Stem export
Yesgithub.com
Supported languages
3 languagesgithub.com
Export formats
MIDIgithub.com
Commercial use
allowedgithub.com

Facts

Core function
SoulX-Singer is a high-fidelity zero-shot singing voice synthesis model for generating realistic voices for unseen singers.github.com · 1 Oct 2026
Pitch and score control
It supports melody-conditioned F0-contour control and score-conditioned MIDI-note control for pitch, rhythm, and expression.github.com · 1 Oct 2026
Voice conversion
SoulX-Singer-SVC converts raw singing audio into a target singer’s voice while preserving melody, rhythm, and lyrics without lyric or MIDI transcriptions.github.com · 1 Oct 2026
Zero-shot operation
The model generates voices for unseen singers without fine-tuning or per-speaker fine-tuning.github.com · 1 Oct 2026
Languages
The system supports Mandarin Chinese, English, and Cantonese.github.com · 1 Oct 2026
Training data
The project reports more than 42,000 hours of aligned vocal, lyric, and note data.arxiv.org · 1 Oct 2026
Editing and cloning
Features include singer-timbre cloning, cross-lingual synthesis, and lyric editing while preserving natural prosody.github.com · 1 Oct 2026
Preprocessing
Its preprocessing toolkit performs vocal separation and dereverberation, F0 extraction, voice activity detection, lyrics transcription, and note transcription.github.com · 1 Oct 2026
MIDI integration
Generated metadata can be exported to MIDI, edited for lyrics, phoneme alignment, pitches, and durations, and imported back for synthesis.github.com · 1 Oct 2026
Web access
The maker provides a SoulX-Singer singing-generation and vocal-conversion demo on Hugging Face Spaces.huggingface.co · 1 Oct 2026
Local deployment
The repository supports local inference through Conda with Python 3.10, pip-installed dependencies, and local WebUI scripts.github.com · 1 Oct 2026
Model distribution
Pretrained synthesis, conversion, and preprocessing models are downloaded through Hugging Face Hub commands.github.com · 1 Oct 2026
License
The code and model weights are released under the Apache 2.0 license for researchers and developers to use.github.com · 1 Oct 2026
Usage restrictions
The maker asks users to respect intellectual property, privacy, and consent and prohibits unauthorized impersonation or deceptive audio.github.com · 1 Oct 2026
Support
The project lists three contact emails and invites technical discussion through WeChat or Soul app groups.github.com · 1 Oct 2026
Purpose
SoulX-Singer is a high-fidelity zero-shot singing voice synthesis model for generating realistic voices for unseen singers.github.com · 1 Oct 2026
Control modes
It supports melody-conditioned F0 contour control and score-conditioned MIDI note control for pitch, rhythm, and expression.github.com · 1 Oct 2026
Dataset scale
The stated training dataset contains more than 42,000 hours of aligned vocals, lyrics, and notes.github.com · 1 Oct 2026
MIDI editing
A MIDI Editor supports editing lyrics, phoneme alignment, note pitches, and durations before inference.github.com · 1 Oct 2026
Online access
Soul-AILab provides a running SoulX-Singer demo on Hugging Face Spaces and a separate running MIDI Editor Space.huggingface.co · 1 Oct 2026
Deployment
The repository can be cloned, installed with Conda and pip, and run locally through Python web UI scripts.github.com · 1 Oct 2026
Notable limitation
The maintainers warn that automatic preprocessing may misalign singing audio with lyrics and notes and recommend manual correction.github.com · 1 Oct 2026
License and safety
The project uses Apache 2.0 and asks users to respect intellectual property, privacy, and consent and avoid unauthorized impersonation or deceptive audio.github.com · 1 Oct 2026

Best SoulX-Singer alternatives

See all 12

Where it ranks on Everything Xiaomi

Is SoulX-Singer yours?

Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.

Sources