Diffusers documentation
Kandinsky 6 VAEs
Kandinsky 6 VAEs
Kandinsky 6 uses a causal 3D K-VAE for video super-resolution and the MMAudio mel-spectrogram VAE, paired with a separate BigVGAN MMAudioVocoder, for synchronized audio generation.
Kandinsky6SRVAE
The causal 3D K-VAE used by Kandinsky6SRPipeline. It processes arbitrarily long videos in bounded-memory segments while reproducing the exact output of a single, non-segmented pass.
import torch
from diffusers import Kandinsky6SRVAE
vae = Kandinsky6SRVAE.from_pretrained(
"kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers", subfolder="vae", torch_dtype=torch.bfloat16
)class diffusers.Kandinsky6SRVAE
< source >( in_channels: int = 3out_channels: int = 3latent_channels: int = 64encoder_block_out_channels: tuple[int, ...] = (16, 128, 256, 512, 1024)decoder_block_out_channels: tuple[int, ...] = (16, 256, 512, 1024, 2048)layers_per_block: int = 2temporal_compression_ratio: int = 4temporal_compression_start_level: int = 1scaling_factor: float = 0.910344004631042 )
Parameters
- in_channels (
int, defaults to3) — Number of pixel channels. - out_channels (
int, defaults to3) — Number of reconstructed pixel channels. - latent_channels (
int, defaults to64) — Number of latent channels. - encoder_block_out_channels (
tuple[int, ...], defaults to(16, 128, 256, 512, 1024)) — Output width of the residual blocks at each encoder level; every level but the last halves the spatial size and doubles the width on the way to the next level. - decoder_block_out_channels (
tuple[int, ...], defaults to(16, 256, 512, 1024, 2048)) — Output width of the residual blocks at each decoder level. - layers_per_block (
int, defaults to2) — Number of residual blocks per encoder level; the decoder uses one more per level. - temporal_compression_ratio (
int, defaults to4) — Temporal compression factor;log2of it consecutive levels also compress time. - temporal_compression_start_level (
int, defaults to1) — First level that compresses time. - scaling_factor (
float, defaults to0.910344) — Scale applied to the latents before they enter the diffusion transformer.
Causal 3D K-VAE used by Kandinsky6SRPipeline to encode and decode video.
Videos are processed in temporal segments of 16 pixel frames (plus the leading frame). The causal convolutions carry their padding state between segments, so the segmentation only bounds peak memory and does not change the result.
encode
< source >( x: torch.Tensorreturn_dict: bool = True )
Parameters
- x (
torch.Tensorof shape(batch_size, channels, num_frames, height, width)) — Pixel video in[-1, 1].num_framesshould be1 + k * temporal_compression_ratio. - return_dict (
bool, defaults toTrue) — Whether to return an AutoencoderKLOutput instead of a plain tuple.
Encode a video into its latent distribution.
decode
< source >( z: torch.Tensorreturn_dict: bool = True )
Decode latents into a video.
forward
< source >( sample: torch.Tensorsample_posterior: bool = Falsereturn_dict: bool = Truegenerator: torch.Generator | None = None ) → ~models.autoencoder_kl.DecoderOutput or tuple
Parameters
- sample (
torch.Tensorof shape(batch_size, channels, num_frames, height, width)) — Pixel video in[-1, 1].num_framesshould be1 + k * temporal_compression_ratio. - sample_posterior (
bool, optional, defaults toFalse) — Whether to sample from the latent posterior instead of using its mode. - return_dict (
bool, optional, defaults toTrue) — Whether to return a~models.autoencoder_kl.DecoderOutputinstead of a plain tuple. - generator (
torch.Generator, optional) — Atorch.Generatorto make sampling deterministic.
Returns
~models.autoencoder_kl.DecoderOutput or tuple
If return_dict is True, a ~models.autoencoder_kl.DecoderOutput is returned, otherwise a plain
tuple is returned. Its sample is the reconstructed video.
MMAudioVAE
The mel-spectrogram VAE used by Kandinsky6TI2VAPipeline when sample_audio=True. Its decode output is a mel
spectrogram; pass it through MMAudioVocoder to get a waveform.
The reference implementation can be found at hkchengrex/MMAudio (MIT license).
import torch
from diffusers import MMAudioVAE
audio_vae = MMAudioVAE.from_pretrained(
"kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers", subfolder="audio_vae", torch_dtype=torch.bfloat16
)class diffusers.MMAudioVAE
< source >( mel_bins: int = 128latent_channels: int = 40hidden_channels: int = 512channel_multipliers: tuple[int, ...] = (1, 2, 4)layers_per_block: int = 2sample_rate: int = 44100n_fft: int = 2048hop_length: int = 512scaling_factor: float = 0.417 )
Parameters
- mel_bins (
int, defaults to128) — Number of mel bins. - latent_channels (
int, defaults to40) — Number of latent channels. - hidden_channels (
int, defaults to512) — Base width of the autoencoder. - channel_multipliers (
tuple[int, ...], defaults to(1, 2, 4)) — Width multipliers of the autoencoder levels. - layers_per_block (
int, defaults to2) — Residual blocks per encoder level; the decoder uses one more per level. - sample_rate (
int, defaults to44100) — Waveform sample rate. - n_fft (
int, defaults to2048) — FFT size of the mel front end. - hop_length (
int, defaults to512) — Hop length of the mel front end. Must match the total upsampling factor of the MMAudioVocoder this VAE is paired with. - scaling_factor (
float, defaults to0.417) — Scale applied to the latents before they enter the diffusion transformer.
Audio VAE of Kandinsky6TI2VAPipeline: a magnitude-preserving autoencoder over log-mel spectrograms (MMAudio, https://arxiv.org/abs/2412.15322).
encode turns a waveform into a latent distribution; decode turns latents back into a mel spectrogram, which MMAudioVocoder then turns into a waveform. One latent frame covers hop_length * 2 samples.
encode
< source >( audio: torch.Tensorreturn_dict: bool = True )
Parameters
- audio (
torch.Tensorof shape(batch_size, num_samples)) — Mono waveform in[-1, 1]atsample_rate. - return_dict (
bool, defaults toTrue) — Whether to return an AutoencoderKLOutput instead of a plain tuple.
Encode a waveform into its latent distribution.
decode
< source >( z: torch.Tensorreturn_dict: bool = True )
Decode latents into a mel spectrogram.
forward
< source >( sample: torch.Tensorsample_posterior: bool = Falsereturn_dict: bool = Truegenerator: torch.Generator | None = None ) → ~models.autoencoder_kl.DecoderOutput or tuple
Parameters
- sample (
torch.Tensorof shape(batch_size, num_samples)) — Mono waveform in[-1, 1]atsample_rateto encode and reconstruct as a mel spectrogram. - sample_posterior (
bool, optional, defaults toFalse) — Whether to sample from the latent posterior instead of using its mode. - return_dict (
bool, optional, defaults toTrue) — Whether to return a~models.autoencoder_kl.DecoderOutputinstead of a plain tuple. - generator (
torch.Generator, optional) — Atorch.Generatorto make sampling deterministic.
Returns
~models.autoencoder_kl.DecoderOutput or tuple
If return_dict is True, a ~models.autoencoder_kl.DecoderOutput is returned, otherwise a plain
tuple is returned. Its sample is the reconstructed mel spectrogram of shape (batch_size, mel_bins, num_mel_frames).
MMAudioVocoder
Adapted from the BigVGAN-v2 vocoder MMAudio bundles, itself from NVIDIA/BigVGAN (MIT license), with the anti-aliased Snake activations of alias-free-torch (Apache License 2.0).
class diffusers.MMAudioVocoder
< source >( num_mels: int = 128upsample_initial_channel: int = 1536upsample_rates: tuple[int, ...] = (8, 4, 2, 2, 2, 2)upsample_kernel_sizes: tuple[int, ...] = (16, 8, 4, 4, 4, 4)resblock_kernel_sizes: tuple[int, ...] = (3, 7, 11)resblock_dilation_sizes: tuple[tuple[int, ...], ...] = ((1, 3, 5), (1, 3, 5), (1, 3, 5)) )
Parameters
- num_mels (
int, defaults to128) — Number of mel bins of the input spectrogram. Must match the paired MMAudioVAE’smel_bins. - upsample_initial_channel (
int, defaults to1536) — Width of the first layer. - upsample_rates (
tuple[int, ...], defaults to(8, 4, 2, 2, 2, 2)) — Upsampling factors of the vocoder stages. Their product is the total upsampling factor and must match the paired MMAudioVAE’shop_length. - upsample_kernel_sizes (
tuple[int, ...], defaults to(16, 8, 4, 4, 4, 4)) — Transposed-convolution kernel sizes of the vocoder stages. - resblock_kernel_sizes (
tuple[int, ...], defaults to(3, 7, 11)) — Kernel sizes of the residual blocks. - resblock_dilation_sizes (
tuple[tuple[int, ...], ...], defaults to((1, 3, 5), (1, 3, 5), (1, 3, 5))) — Dilations of the residual blocks.
BigVGAN-v2 vocoder (https://github.com/NVIDIA/BigVGAN, MIT license) with the anti-aliased Snake activations of https://github.com/junjun3518/alias-free-torch (Apache 2.0), turning the mel spectrograms MMAudioVAE decodes into waveforms for Kandinsky6TI2VAPipeline.
forward
< source >( mel: torch.Tensorreturn_dict: bool = True )
Parameters
- mel (
torch.Tensorof shape(batch_size, num_mels, num_mel_frames)) — Mel spectrogram, as decoded by MMAudioVAE. - return_dict (
bool, defaults toTrue) — Whether to return a~models.autoencoder_kl.DecoderOutputinstead of a plain tuple.