Skip to content

Optimization catalog

Each optimization has one short page with executable code, support boundaries, research references, and upstream GitHub source.

Generated by scripts/generate_optimization_pages.py.

Optimization passes

Optimization Availability Paper GitHub
Torch compile Registered public pass: compile Available PyTorch
FlashAttention-4 Registered public pass: flash-attention-4 Available FlashAttention
Custom kernels Registered public pass: custom-kernels Available Triton
Codec kernels Registered public pass: codec-kernels Available Triton; NVIDIA CUTLASS / CuTe
Diffusion cache Registered public pass: diffusion-cache Available DeepCache; TeaCache
Diffusion sampling Registered public pass: diffusion-sampling Available DPM-Solver

Optional source backends

Optimization Availability Paper GitHub
HQQ Optional library; no registered VoiceHub pass Available HQQ
GemLite Optional library; no registered VoiceHub pass Available GemLite
audio.cpp External runtime; no registered VoiceHub pass No dedicated paper audio.cpp

Serving backends

Optimization Availability Paper GitHub
vLLM Built-in VoiceHub HTTP client; the capability registry controls verified model pairs Available vLLM; vLLM-Omni
SGLang Built-in VoiceHub HTTP client; the capability registry controls verified model pairs Available SGLang; SGLang-Omni