vllm_omni.diffusion.models.ltx2.ltx2_sequence_parallel ¶
LTX2VideoToAudioParallelAttention ¶
LTX video-to-audio SP with replicated audio queries.
Video K/V arrive sequence-sharded, while the much shorter audio query is replicated. Redistribute only K/V across sequence/head dimensions, slice audio query heads locally, then gather output heads. This uses three collectives instead of standard Ulysses' four and keeps audio replicated.
post_attention ¶
post_attention(
attn_output: Tensor,
ctx: ParallelAttentionContext | None,
) -> Tensor
pre_attention ¶
pre_attention(
query: Tensor,
key: Tensor,
value: Tensor,
attn_metadata: AttentionMetadata | None,
)