Zum Inhalt

TTS capabilities and adaptation boundaries

This matrix describes the current VoiceHub provider contract. Diffusion includes flow-matching models. Task codes are C voice cloning from a reference/profile/embedding, V text-guided voice design, D dialogue, and S explicit style, emotion, or description control. Language abbreviations follow the registered default checkpoint and its release card.

Model type Family Verified languages C/V/D/S Fine-tuning and LoRA boundary
orpheustts LLM en S Raw text/audio or SNAC codes train the LM; SNAC stays frozen. LoRA: no.
dia LLM en D Raw text/audio trains the Dia-1.6B-0626 encoder/decoder; native DAC prepares targets. The legacy checkpoint is unsupported. LoRA: no.
vui LLM en C (codec prompt) Prepared text IDs and Fluac codes train a reconstructed LM objective; Fluac stays frozen. LoRA: no.
chatterbox Hybrid ar, da, de, el, en, es, fi, fr, he, hi, it, ja, ko, ms, nl, no, pl, pt, ru, sv, sw, tr, zh C, S Raw or prepared data trains T3 LM and S3Gen flow in separate jobs, not one joint job. LoRA: T3 LM only.
kokoro Acoustic en-US, en-GB, es, fr, hi, it, pt-BR, ja, zh S Prepared partial PL-BERT, duration/prosody, and decoder objectives; not the raw-data/full author recipe. LoRA: no.
echo Diffusion en C Prepared Fish-codec latents train an experimental reconstructed DiT-flow objective; codec stays frozen. LoRA: no.
conversationtts LLM en, zh, yue C, D Raw data or Mimi codes train the language/depth model; Mimi stays frozen. LoRA: no.
llasa LLM zh, en, de, fr, ja, ko, nl, es, it, pt, pl C Raw data or XCodec2 codes train only the LM; XCodec2 stays frozen. LoRA: no.
cosyvoice Hybrid zh, en, ja, ko, de, es, fr, it, ru C, S Raw/prepared data trains LM, flow, HiFT generator, or discriminator in separate jobs; S3Tokenizer stays frozen and CAMPPlus embeddings remain explicit. LoRA: no.
f5tts Diffusion en, zh C Waveform/mel plus prepared text IDs train the full DiT flow; Vocos stays frozen. Chinese needs explicit pinyin or a native normalizer. LoRA: no.
gptsovits Hybrid zh, en, ja, ko, yue C Prepared staged S1, S2-generator, and S2-discriminator jobs cover V1/V2/V2Pro/V2ProPlus. Korean and Cantonese start with V2; V3, V4, and LoRA are unsupported.
melotts VITS en, fr, ja, es, zh, ko — Prepared linguistic/BERT/spectrogram/audio tensors train full VITS/GAN phases; raw multilingual preparation and author-resumable state are not provided. LoRA: no.
openvoice VITS en, es, fr, zh, ja, ko C Opt-in reconstructed paired-waveform converter training only; no released upstream recipe or quality-parity claim. LoRA: no.
outetts LLM en, ar, zh, nl, fr, de, it, ja, ko, lt, ru, es, pt, be, bn, ka, hu, lv, fa, pl, sw, ta, uk C (profile) Prepared V3 profiles or token labels train the full LM; DAC stays frozen and raw audio is unsupported. LoRA: no.
parlertts LLM en S Raw description/text/audio or DAC codes train the decoder and T5 by default; T5 may be frozen and DAC always is. LoRA: no.
styletts2 Diffusion en-US C, S Prepared reconstructed generator/GAN phases; no raw G2P/alignment, WavLM objective, or author-resume parity. LoRA: no.
mosstts LLM zh, yue, en, ar, cs, da, de, nl, es, fr, fi, el, he, hi, hu, ja, it, ko, mk, ms, ru, fa, pl, pt, sv, ro, sw, tl, th, tr, vi C Raw audio or RVQ codes train the complete semantic graph for each supported variant; its codec stays frozen. LoRA: no.
qwen3tts LLM zh, en, ja, ko, de, fr, ru, pt, es, it C, V, S Base checkpoints only: prepared 16-codebook SFT trains talker/residual predictor fully by default, or native LoRA adapts their attention/MLP projections. Speaker encoder stays frozen; CustomVoice/VoiceDesign are not training starts.
irodoritts Diffusion ja C, V, S Raw audio or latents train RF-DiT and optional duration prediction; Semantic-DACVAE stays frozen. LoRA: no.
zonos LLM en, ja, zh, fr, de C, S Raw/codes with explicit or injected phonemes train a reconstructed dense-Transformer LM objective; DAC stays frozen. Mamba-2 is unsupported. LoRA: no.
zonos2 LLM en, zh, ja, ko, ru, it, pt, fr, es, vi, de, he, nl, sv, hi, ta, te, th, no, bn, tl, ar, da, id, pl, uk, ro, fi, hu, lt, et, sk, hr, lv C, S Raw audio or codes train the dense/MoE acoustic LM under a reconstructed objective; codec is outside the trainable graph. LoRA: no.
voxcpm Hybrid zh, en, ar, my, da, nl, fi, fr, de, el, he, hi, id, it, ja, km, ko, lo, ms, no, pl, pt, ru, es, sw, sv, tl, th, tr, vi C, V, S Raw audio or latents train the MiniCPM/flow graph, or native LoRA targets LM, DiT, and optional projections. AudioVAE stays frozen.
omnivoice LLM aae, aal, aao, ab, abb, abn, abr, abs, abv, acm, acw, acx, adf, adx, ady, aeb, aec, af, afb, afo, ahl, ahs, ajg, aju, ala, aln, alo, am, amu, an, anc, ank, anp, anw, aom, apc, apd, arb, arq, ars, ary, arz, as, ast, avl, awo, ayl, ayp, az, ba, bag, bas, bax, bba, bbj, bbl, bbu, bce, bci, bcs, bcy, bda, bde, bdm, be, beb, bew, bfd, bft, bg, bgp, bhb, bhh, bho, bhp, bhr, bjj, bjk, bjn, bjt, bkh, bkm, bky, bmm, bmq, bn, bnm, bnn, bns, bo, bou, bqg, br, bra, brh, bri, brx, bs, bsh, bsj, bsk, btm, btv, bug, bum, buo, bux, bwr, bxf, byc, bys, byv, byx, bzc, bzw, ca, ccg, ceb, cen, cfa, cgg, chq, cjk, ckb, ckl, ckr, cky, cnh, cpy, cs, cte, ctl, cut, cux, cv, cy, da, dag, dar, dav, dbd, dcc, de, deg, dgh, dgo, dje, dmk, dml, dru, dty, dua, dv, dyu, dzg, ebr, ebu, ego, eiv, eko, ekr, el, elm, en, eo, es, esu, et, eto, ets, etu, eu, ewo, ext, eyo, fa, fan, fat, ff, ffm, fi, fia, fil, fip, fkk, fmp, fr, fub, fuc, fue, fuf, fuh, fui, fuq, fuv, fy, ga, gbm, gbr, gby, gcc, gdf, gej, ges, ggg, gid, gig, giz, gjk, gju, gl, glw, gn, gol, gom, gsl, gu, gui, gur, guz, gv, gwc, gwe, gwt, gya, gyz, ha, hah, hao, haw, haz, hbb, he, hem, hi, hia, hkk, hla, hno, hoj, hr, hsb, ht, hu, hue, hul, hux, hwo, hy, hz, ia, ibb, id, ida, idu, ig, ijc, ijn, ik, ikw, is, ish, iso, it, its, itw, itz, ja, jal, jax, jgo, jmx, jns, jqr, juk, juo, jv, ka, kab, kai, kaj, kam, kbd, kbl, kbt, kcq, kdh, kea, keu, kfe, kfk, kfp, khg, khw, kj, kjc, kjk, kk, kln, kls, km, kmr, kmy, kn, kna, knn, ko, kol, koo, kpo, kqo, ks, ksd, ksf, kto, kuh, kvx, kw, kwm, kxp, ky, kyx, lag, lb, lcm, ldb, lg, lij, lir, lkb, lla, ln, lnu, lo, loa, lrk, lss, lt, ltg, lto, lua, luo, lus, lv, lwg, mab, maf, mai, mau, max, mbo, mcf, mcn, mcx, mdd, mde, mdf, mek, mer, meu, mfm, mfn, mfo, mfv, mgg, mgi, mhk, mhr, mi, mig, miu, mk, mkf, mki, ml, mlq, mn, mne, mni, mqy, mr, mrj, mrr, mrt, ms, mse, msh, msw, mt, mtr, mtu, mtx, mua, mug, mui, mve, mvy, mxs, mxu, mxy, my, myv, mzl, nal, nan, nap, nb, nbh, ncf, nco, ncx, ndi, ng, ngi, nhg, nhi, nhn, nhq, nja, nl, nla, nlv, nmg, nmz, nn, nnh, no, noe, npi, nso, ny, nyu, oc, odk, odu, ogo, om, orc, oru, ory, os, pa, pbs, pbt, pbu, pcm, pex, phl, phr, pip, piy, pko, pl, plk, plt, pmq, pms, pmy, pnb, poc, poe, pow, prq, ps, pst, pt, pua, pwn, qug, qum, qup, qur, qus, quv, qux, quy, qva, qvi, qvj, qvl, qwa, qws, qxa, qxp, qxt, qxu, qxw, rag, rm, ro, rob, rof, roo, rth, ru, rup, rw, sa, sah, sat, sau, say, sbn, sc, scl, scn, sd, sei, shu, si, sip, siw, sjr, sk, skg, skr, sl, sn, snc, snk, so, sol, sps, sq, sr, src, sro, ssi, ste, sua, sv, sva, sw, szy, ta, tan, tar, tay, tbf, tcf, tcy, tdn, tdx, te, tg, tgc, th, the, thq, thr, thv, ti, tig, tio, tk, tkg, tkt, tli, tlp, tn, tok, tpl, tpz, tqp, tr, trp, trq, trv, trw, tt, ttj, ttr, ttu, tui, tul, tuq, tuv, tuy, tvo, tvu, tw, twu, txs, txy, udl, ug, uk, uki, umb, ur, ush, uz, uzn, vai, var, ver, vi, vmc, vmj, vmm, vmp, vmz, vot, vro, wbl, wci, weo, wes, wja, wji, wo, wof, xh, xhe, xka, xmf, xmv, xmw, xpe, xti, xtu, yaq, yav, yay, ydd, ydg, yer, yes, yi, yo, yue, zga, zgh, zh, zoc, zoh, zor, zpv, zpy, ztg, ztn, ztp, zts, ztu, zu, zza C, V Raw audio or codes train the complete masked-token model; Higgs Audio codec stays frozen. LoRA: no.
higgstts LLM en, zh, de, ko C, S Raw audio or codes train the dual-FFN decoder; HuBERT/DAC tokenizer stays frozen and VoiceHub owns the unpublished optimizer schedule. LoRA: no.
xtts LLM en, es, fr, de, it, pt, pl, tr, ru, nl, cs, ar, zh-CN, hu, ko, ja, hi C Raw audio or codes train only the GPT; DVAE, speaker encoder, and HiFi-GAN stay frozen. No vocoder/GAN phase or LoRA.
vibevoice Hybrid en —* Prepared latents train the non-streaming 1.5B LM/connectors/diffusion head; codecs stay frozen. Realtime/default-checkpoint FT and high-level TTS fail closed. LoRA: no.
fishtts LLM zh, en, ja, ko, es, pt, ar, ru, fr, de, sv, it, tr, no, nl, cy, eu, ca, da, gl, ta, hu, fi, pl, et, hi, la, ur, th, vi, jw, bn, yo, sl, cs, sw, nn, he, ms, uk, id, kk, bg, lv, my, tl, sk, ne, fa, af, el, bo, hr, ro, sn, mi, yi, am, be, km, is, az, sd, br, sq, ps, mn, ht, ml, sr, sa, te, ka, bs, pa, lt, kn, si, hy, mr, as, gu, fo C Prepared tokens train the slow and fast semantic model only; ModifiedDAC stays frozen. Provider LoRA is rejected; derivatives are non-commercial.
csm LLM en C, D Raw conversation/audio or Mimi codes train the CSM backbone/depth decoder; Mimi stays frozen. LoRA: no.
neutts LLM en C, S NeuTTS-Air only: raw audio or codes train the LM while NeuCodec stays frozen. Nano, multilingual Nano, and default 2E FT fail closed. LoRA: no.
supertonic Diffusion en, ko, ja, ar, bg, cs, da, de, el, es, et, fi, fr, hi, hr, hu, id, it, lt, lv, nl, pl, pt, ro, ru, sk, sl, sv, tr, uk, vi S Prepared style, duration, and latent targets train reconstructed graph losses; no raw-data or complete author recipe. LoRA: no.
inflecttts VITS en-US — Prepared phonemes/spectrogram/audio train a full-VITS warm start with newly initialized posterior/discriminator; it is not author-resumable. LoRA: no.
bark LLM de, en, es, fr, hi, it, ja, ko, pl, pt, ru, tr, zh S Prepared tokens train semantic, coarse, or fine stages separately; Encodec stays frozen and no joint raw-audio recipe is provided. LoRA: no.
speecht5 Acoustic en C (x-vector) Raw text/audio trains the complete spectrogram model; HiFi-GAN stays frozen. LoRA: no.
vits VITS en — Full raw-waveform adversarial FT requires an explicit checkpoint acoustic config; the generator-only prepared route is a partial warm start. LoRA: no.

* VibeVoice exposes verified low-level realtime stages, but its unified high-level waveform-generation contract is intentionally unavailable.

See TTS training support for the detailed data, checkpoint, objective, and export contracts.