UTAI SYNTHESIZER / SINGING-VOICE SYNTHESIS DAW
Touch the staff to place notes · sweep the cursor to play them

UTAI = 歌 (uta) × AI / a singing-synthesis DAW

From a dozen minutes of singing,
comes your very own 「utahime」.

Sing freely, with the power of AI.

you are visitor no. 000142
— Hz

WHAT IT IS

A singing DAW that plays voice-conversion models like virtual singers.

UtaiSynthesizer turns high-freedom SVC tuning into a full singing-synthesis workflow — record a little, train a little, and it sings your score.

THREE EDGES

Three things that set it apart.

  • High-freedom SVC tuning

    Dual backend — RVC for speed, SoVITS for quality — with shallow diffusion and voice blending (speaker-embedding interpolation).

  • Score2ConVec — SVC as SVS

    Our own model turns a score into ContentVec features, so a voice-conversion model sings from notation. A dozen minutes of dry vocals plus an hour or two of training — no UTAU slicing, no hours of finely-labelled data.

  • A complete DAW

    Piano roll, multi-track, node workflow, vocal separation (MSST), a training monitor, 7-language G2P, range extension — and export to audio · ust · ustx · midi.

HOW IT WORKS

From notation to voice, in one line.

A clean, one-way pipeline — no tangled node graph.

SOURCE

Open source, in the open.

Two repos, out in the open: the synthesizer itself, and Score2ConVec — the model that makes it sing.