vllm_omni.diffusion.models.hunyuan_image3.layers.nvidia ¶
NVIDIA (CUDA) implementations.
Modules:
| Name | Description |
|---|---|
autoencoder_blocks | ResnetBlock for the HunyuanImage3 autoencoder — NVIDIA CUDA + Triton implementation. |
transformer_blocks |
|
ResBlock ¶
Bases: Module
A residual block that can optionally change the number of channels. Args: in_channels (int): The number of input channels. emb_channels (int): The number of timestep embedding channels. dropout (float): The rate of dropout. out_channels (int, optional): If specified, the number of output channels. use_conv (bool, optional): If True and out_channels is specified, use a spatial convolution instead of a smaller 1x1 convolution to change the channels in the skip connection. dims (int, optional): Determines if the signal is 1D, 2D, or 3D. up (bool, optional): If True, use this block for upsampling. down (bool, optional): If True, use this block for downsampling.
emb_layers instance-attribute ¶
emb_layers = nn.Sequential(
nn.SiLU(),
nn.Linear(
emb_channels,
2 * self.out_channels,
**factory_kwargs,
),
)
in_layers instance-attribute ¶
in_layers = nn.Sequential(
normalization(self.in_channels, **factory_kwargs),
nn.SiLU(),
conv_nd(
dims,
self.in_channels,
self.out_channels,
3,
padding=1,
**factory_kwargs,
),
)
out_layers instance-attribute ¶
out_layers = nn.Sequential(
normalization(self.out_channels, **factory_kwargs),
nn.SiLU(),
nn.Dropout(p=dropout),
zero_module(
conv_nd(
dims,
self.out_channels,
self.out_channels,
3,
padding=1,
**factory_kwargs,
)
),
)