TTS capabilities and adaptation boundaries¶
This matrix describes the current VoiceHub provider contract. Diffusion includes flow-matching models. Task codes are C voice cloning from a reference/profile/embedding, V text-guided voice design, D dialogue, and S explicit style, emotion, or description control. Language abbreviations follow the registered default checkpoint and its release card.
| Model type | Family | Verified languages | C/V/D/S | Fine-tuning and LoRA boundary |
|---|---|---|---|---|
orpheustts |
LLM | en | S | Raw text/audio or SNAC codes train the LM; SNAC stays frozen. LoRA: no. |
dia |
LLM | en | D | Raw text/audio trains the Dia-1.6B-0626 encoder/decoder; native DAC prepares targets. The legacy checkpoint is unsupported. LoRA: no. |
vui |
LLM | en | C (codec prompt) | Prepared text IDs and Fluac codes train a reconstructed LM objective; Fluac stays frozen. LoRA: no. |
chatterbox |
Hybrid | ar, da, de, el, en, es, fi, fr, he, hi, it, ja, ko, ms, nl, no, pl, pt, ru, sv, sw, tr, zh | C, S | Raw or prepared data trains T3 LM and S3Gen flow in separate jobs, not one joint job. LoRA: T3 LM only. |
kokoro |
Acoustic | en-US, en-GB, es, fr, hi, it, pt-BR, ja, zh | S | Prepared partial PL-BERT, duration/prosody, and decoder objectives; not the raw-data/full author recipe. LoRA: no. |
echo |
Diffusion | en | C | Prepared Fish-codec latents train an experimental reconstructed DiT-flow objective; codec stays frozen. LoRA: no. |
conversationtts |
LLM | en, zh, yue | C, D | Raw data or Mimi codes train the language/depth model; Mimi stays frozen. LoRA: no. |
llasa |
LLM | zh, en, de, fr, ja, ko, nl, es, it, pt, pl | C | Raw data or XCodec2 codes train only the LM; XCodec2 stays frozen. LoRA: no. |
cosyvoice |
Hybrid | zh, en, ja, ko, de, es, fr, it, ru | C, S | Raw/prepared data trains LM, flow, HiFT generator, or discriminator in separate jobs; S3Tokenizer stays frozen and CAMPPlus embeddings remain explicit. LoRA: no. |
f5tts |
Diffusion | en, zh | C | Waveform/mel plus prepared text IDs train the full DiT flow; Vocos stays frozen. Chinese needs explicit pinyin or a native normalizer. LoRA: no. |
gptsovits |
Hybrid | zh, en, ja, ko, yue | C | Prepared staged S1, S2-generator, and S2-discriminator jobs cover V1/V2/V2Pro/V2ProPlus. Korean and Cantonese start with V2; V3, V4, and LoRA are unsupported. |
melotts |
VITS | en, fr, ja, es, zh, ko | — | Prepared linguistic/BERT/spectrogram/audio tensors train full VITS/GAN phases; raw multilingual preparation and author-resumable state are not provided. LoRA: no. |
openvoice |
VITS | en, es, fr, zh, ja, ko | C | Opt-in reconstructed paired-waveform converter training only; no released upstream recipe or quality-parity claim. LoRA: no. |
outetts |
LLM | en, ar, zh, nl, fr, de, it, ja, ko, lt, ru, es, pt, be, bn, ka, hu, lv, fa, pl, sw, ta, uk | C (profile) | Prepared V3 profiles or token labels train the full LM; DAC stays frozen and raw audio is unsupported. LoRA: no. |
parlertts |
LLM | en | S | Raw description/text/audio or DAC codes train the decoder and T5 by default; T5 may be frozen and DAC always is. LoRA: no. |
styletts2 |
Diffusion | en-US | C, S | Prepared reconstructed generator/GAN phases; no raw G2P/alignment, WavLM objective, or author-resume parity. LoRA: no. |
mosstts |
LLM | zh, yue, en, ar, cs, da, de, nl, es, fr, fi, el, he, hi, hu, ja, it, ko, mk, ms, ru, fa, pl, pt, sv, ro, sw, tl, th, tr, vi | C | Raw audio or RVQ codes train the complete semantic graph for each supported variant; its codec stays frozen. LoRA: no. |
qwen3tts |
LLM | zh, en, ja, ko, de, fr, ru, pt, es, it | C, V, S | Base checkpoints only: prepared 16-codebook SFT trains talker/residual predictor fully by default, or native LoRA adapts their attention/MLP projections. Speaker encoder stays frozen; CustomVoice/VoiceDesign are not training starts. |
irodoritts |
Diffusion | ja | C, V, S | Raw audio or latents train RF-DiT and optional duration prediction; Semantic-DACVAE stays frozen. LoRA: no. |
zonos |
LLM | en, ja, zh, fr, de | C, S | Raw/codes with explicit or injected phonemes train a reconstructed dense-Transformer LM objective; DAC stays frozen. Mamba-2 is unsupported. LoRA: no. |
zonos2 |
LLM | en, zh, ja, ko, ru, it, pt, fr, es, vi, de, he, nl, sv, hi, ta, te, th, no, bn, tl, ar, da, id, pl, uk, ro, fi, hu, lt, et, sk, hr, lv | C, S | Raw audio or codes train the dense/MoE acoustic LM under a reconstructed objective; codec is outside the trainable graph. LoRA: no. |
voxcpm |
Hybrid | zh, en, ar, my, da, nl, fi, fr, de, el, he, hi, id, it, ja, km, ko, lo, ms, no, pl, pt, ru, es, sw, sv, tl, th, tr, vi | C, V, S | Raw audio or latents train the MiniCPM/flow graph, or native LoRA targets LM, DiT, and optional projections. AudioVAE stays frozen. |
omnivoice |
LLM | aae, aal, aao, ab, abb, abn, abr, abs, abv, acm, acw, acx, adf, adx, ady, aeb, aec, af, afb, afo, ahl, ahs, ajg, aju, ala, aln, alo, am, amu, an, anc, ank, anp, anw, aom, apc, apd, arb, arq, ars, ary, arz, as, ast, avl, awo, ayl, ayp, az, ba, bag, bas, bax, bba, bbj, bbl, bbu, bce, bci, bcs, bcy, bda, bde, bdm, be, beb, bew, bfd, bft, bg, bgp, bhb, bhh, bho, bhp, bhr, bjj, bjk, bjn, bjt, bkh, bkm, bky, bmm, bmq, bn, bnm, bnn, bns, bo, bou, bqg, br, bra, brh, bri, brx, bs, bsh, bsj, bsk, btm, btv, bug, bum, buo, bux, bwr, bxf, byc, bys, byv, byx, bzc, bzw, ca, ccg, ceb, cen, cfa, cgg, chq, cjk, ckb, ckl, ckr, cky, cnh, cpy, cs, cte, ctl, cut, cux, cv, cy, da, dag, dar, dav, dbd, dcc, de, deg, dgh, dgo, dje, dmk, dml, dru, dty, dua, dv, dyu, dzg, ebr, ebu, ego, eiv, eko, ekr, el, elm, en, eo, es, esu, et, eto, ets, etu, eu, ewo, ext, eyo, fa, fan, fat, ff, ffm, fi, fia, fil, fip, fkk, fmp, fr, fub, fuc, fue, fuf, fuh, fui, fuq, fuv, fy, ga, gbm, gbr, gby, gcc, gdf, gej, ges, ggg, gid, gig, giz, gjk, gju, gl, glw, gn, gol, gom, gsl, gu, gui, gur, guz, gv, gwc, gwe, gwt, gya, gyz, ha, hah, hao, haw, haz, hbb, he, hem, hi, hia, hkk, hla, hno, hoj, hr, hsb, ht, hu, hue, hul, hux, hwo, hy, hz, ia, ibb, id, ida, idu, ig, ijc, ijn, ik, ikw, is, ish, iso, it, its, itw, itz, ja, jal, jax, jgo, jmx, jns, jqr, juk, juo, jv, ka, kab, kai, kaj, kam, kbd, kbl, kbt, kcq, kdh, kea, keu, kfe, kfk, kfp, khg, khw, kj, kjc, kjk, kk, kln, kls, km, kmr, kmy, kn, kna, knn, ko, kol, koo, kpo, kqo, ks, ksd, ksf, kto, kuh, kvx, kw, kwm, kxp, ky, kyx, lag, lb, lcm, ldb, lg, lij, lir, lkb, lla, ln, lnu, lo, loa, lrk, lss, lt, ltg, lto, lua, luo, lus, lv, lwg, mab, maf, mai, mau, max, mbo, mcf, mcn, mcx, mdd, mde, mdf, mek, mer, meu, mfm, mfn, mfo, mfv, mgg, mgi, mhk, mhr, mi, mig, miu, mk, mkf, mki, ml, mlq, mn, mne, mni, mqy, mr, mrj, mrr, mrt, ms, mse, msh, msw, mt, mtr, mtu, mtx, mua, mug, mui, mve, mvy, mxs, mxu, mxy, my, myv, mzl, nal, nan, nap, nb, nbh, ncf, nco, ncx, ndi, ng, ngi, nhg, nhi, nhn, nhq, nja, nl, nla, nlv, nmg, nmz, nn, nnh, no, noe, npi, nso, ny, nyu, oc, odk, odu, ogo, om, orc, oru, ory, os, pa, pbs, pbt, pbu, pcm, pex, phl, phr, pip, piy, pko, pl, plk, plt, pmq, pms, pmy, pnb, poc, poe, pow, prq, ps, pst, pt, pua, pwn, qug, qum, qup, qur, qus, quv, qux, quy, qva, qvi, qvj, qvl, qwa, qws, qxa, qxp, qxt, qxu, qxw, rag, rm, ro, rob, rof, roo, rth, ru, rup, rw, sa, sah, sat, sau, say, sbn, sc, scl, scn, sd, sei, shu, si, sip, siw, sjr, sk, skg, skr, sl, sn, snc, snk, so, sol, sps, sq, sr, src, sro, ssi, ste, sua, sv, sva, sw, szy, ta, tan, tar, tay, tbf, tcf, tcy, tdn, tdx, te, tg, tgc, th, the, thq, thr, thv, ti, tig, tio, tk, tkg, tkt, tli, tlp, tn, tok, tpl, tpz, tqp, tr, trp, trq, trv, trw, tt, ttj, ttr, ttu, tui, tul, tuq, tuv, tuy, tvo, tvu, tw, twu, txs, txy, udl, ug, uk, uki, umb, ur, ush, uz, uzn, vai, var, ver, vi, vmc, vmj, vmm, vmp, vmz, vot, vro, wbl, wci, weo, wes, wja, wji, wo, wof, xh, xhe, xka, xmf, xmv, xmw, xpe, xti, xtu, yaq, yav, yay, ydd, ydg, yer, yes, yi, yo, yue, zga, zgh, zh, zoc, zoh, zor, zpv, zpy, ztg, ztn, ztp, zts, ztu, zu, zza | C, V | Raw audio or codes train the complete masked-token model; Higgs Audio codec stays frozen. LoRA: no. |
higgstts |
LLM | en, zh, de, ko | C, S | Raw audio or codes train the dual-FFN decoder; HuBERT/DAC tokenizer stays frozen and VoiceHub owns the unpublished optimizer schedule. LoRA: no. |
xtts |
LLM | en, es, fr, de, it, pt, pl, tr, ru, nl, cs, ar, zh-CN, hu, ko, ja, hi | C | Raw audio or codes train only the GPT; DVAE, speaker encoder, and HiFi-GAN stay frozen. No vocoder/GAN phase or LoRA. |
vibevoice |
Hybrid | en | —* | Prepared latents train the non-streaming 1.5B LM/connectors/diffusion head; codecs stay frozen. Realtime/default-checkpoint FT and high-level TTS fail closed. LoRA: no. |
fishtts |
LLM | zh, en, ja, ko, es, pt, ar, ru, fr, de, sv, it, tr, no, nl, cy, eu, ca, da, gl, ta, hu, fi, pl, et, hi, la, ur, th, vi, jw, bn, yo, sl, cs, sw, nn, he, ms, uk, id, kk, bg, lv, my, tl, sk, ne, fa, af, el, bo, hr, ro, sn, mi, yi, am, be, km, is, az, sd, br, sq, ps, mn, ht, ml, sr, sa, te, ka, bs, pa, lt, kn, si, hy, mr, as, gu, fo | C | Prepared tokens train the slow and fast semantic model only; ModifiedDAC stays frozen. Provider LoRA is rejected; derivatives are non-commercial. |
csm |
LLM | en | C, D | Raw conversation/audio or Mimi codes train the CSM backbone/depth decoder; Mimi stays frozen. LoRA: no. |
neutts |
LLM | en | C, S | NeuTTS-Air only: raw audio or codes train the LM while NeuCodec stays frozen. Nano, multilingual Nano, and default 2E FT fail closed. LoRA: no. |
supertonic |
Diffusion | en, ko, ja, ar, bg, cs, da, de, el, es, et, fi, fr, hi, hr, hu, id, it, lt, lv, nl, pl, pt, ro, ru, sk, sl, sv, tr, uk, vi | S | Prepared style, duration, and latent targets train reconstructed graph losses; no raw-data or complete author recipe. LoRA: no. |
inflecttts |
VITS | en-US | — | Prepared phonemes/spectrogram/audio train a full-VITS warm start with newly initialized posterior/discriminator; it is not author-resumable. LoRA: no. |
bark |
LLM | de, en, es, fr, hi, it, ja, ko, pl, pt, ru, tr, zh | S | Prepared tokens train semantic, coarse, or fine stages separately; Encodec stays frozen and no joint raw-audio recipe is provided. LoRA: no. |
speecht5 |
Acoustic | en | C (x-vector) | Raw text/audio trains the complete spectrogram model; HiFi-GAN stays frozen. LoRA: no. |
vits |
VITS | en | — | Full raw-waveform adversarial FT requires an explicit checkpoint acoustic config; the generator-only prepared route is a partial warm start. LoRA: no. |
* VibeVoice exposes verified low-level realtime stages, but its unified high-level waveform-generation contract is intentionally unavailable.
See TTS training support for the detailed data, checkpoint, objective, and export contracts.