vllm_omni.core.sched.utils ¶
Shared utilities for omni schedulers.
free_kv_blocks_in_physical_order ¶
Return non-cached request blocks in increasing physical order.
vLLM's block pool deliberately honors the eviction-priority order passed to free_blocks. For a non-caching P/D pool there is no LRU priority to preserve, so physical order lets native connectors coalesce adjacent source/destination pages on the next allocation.
omni_routed_experts_for_request ¶
omni_routed_experts_for_request(
routed_experts: RoutedExpertsLists, request
) -> ndarray | None
Extract per-request routed experts from RoutedExpertsLists using slot_mapping.
Matches upstream RoutedExpertsManager.get() pattern — filters routing_data rows whose slot_mapping entries belong to this request's block_table.