vllm_omni.diffusion.diffusion_kv.initialization ¶
build_native_kv_cache_configs ¶
build_native_kv_cache_configs(
vllm_config: VllmConfig,
worker_specs: list[dict[str, KVCacheSpec]],
available_memory: list[int],
*,
indexes_kv_by_block_stride: bool = False,
) -> tuple[list[KVCacheConfig], KVCacheConfig]
Build rank-local and Scheduler-native configs using vLLM utilities.
indexes_kv_by_block_stride carries the native backend requirement that 0.29 removed from AttentionSpec. Every diffusion backend in tree declares False, so the default reproduces current behaviour; a caller that learns otherwise from its workers passes True here and the block-outermost layout propagates to every rank-local config.
initialize_diffusion_kv_control_plane ¶
initialize_diffusion_kv_control_plane(
executor: DiffusionExecutor,
od_config: OmniDiffusionConfig,
*,
profile_requests: list[OmniDiffusionRequest]
| None = None,
device: device | None = None,
) -> tuple[KVCacheConfig, int, int, VllmConfig] | None
Run the native Worker-spec to Scheduler-config initialization chain.