Realtime AR-Diffusion sessions¶
Status: experimental. This document describes the internal execution contract implemented under
vllm_omni.experimental.ar_diffusion. It does not define a public HTTP or WebSocket schema.
Scope¶
Realtime AR-Diffusion serves world models that generate one autoregressive latent block per request while retaining model state across requests. The generic runtime owns identity, event ordering, session lifecycle, routing, and KV allocation. A model adapter owns model-specific controls and conditioning.
LingBot World v2 is the first single-KV-branch integration. DreamZero uses the same runner capability with two logical CFG branches.
Identities¶
Four identifiers have distinct lifetimes:
| Identifier | Lifetime | Owner | Contract |
|---|---|---|---|
session_id | One persistent world | Session manager | Selects runner KV and model-owned state. |
event_id | One control/prompt update | Transport adapter | Unique and monotonically increasing within a session, including across reset. |
chunk_index | One committed AR block position | Session | Contiguous from zero; reset starts again at zero. |
request_id | One chunk execution | Session | Correlates the tick with its returned metadata. |
AsyncOmni keeps its normal UUID-suffixed engine routing ID. The tick request_id remains inside the immutable AR-Diffusion snapshot and the standard output metadata, while caller-visible OmniRequestOutput.request_id continues to use AsyncOmni's existing external-ID contract. Protocol identity therefore does not require a private entrypoint override.
Request lifecycle¶
ARDiffusionSession.accept_event()validates and queues an immutableARDiffusionSessionEvent. Prompt and opaque control tracks may be updated in the same event.ARDiffusionSession.next_chunk()serializes chunk execution and calls_build_tick_locked(). The resultingARDiffusionTickRequestis a snapshot of the current prompt, reduced controls, ordered event IDs, and chunk index.ARDiffusionOmniTickConsumer.execute_tick()copies the stage sampling parameters, stores the typed tick insampling_params.extra_args["ar_diffusion_tick"], constructs the normal Omni prompt through itsprompt_provider, and invokesAsyncOmni.generate().AsyncOmniallocates its normal unique engine routing ID and submits oneOmniDiffusionRequestthrough the orchestrator.ARDiffusionModelRunner._request_session()parses the typed tick._get_or_create_session()obtains runner-owned paged KV, andbind_ar_diffusion_state()exposes it to the pipeline only for the duration ofexecute_model().- The model pipeline validates the tick and generates exactly one latent block. LingBot interprets camera controls in
LingBotCameraControlReducerandLingBotWorldCausalDMDPipeline.forward(); the generic runtime does not parse camera or action schemas. - The pipeline returns one
DiffusionOutputwith the generated payload and anARDiffusionChunkMetadatamapping undermetadata["ar_diffusion"].AsyncOmniexposes it asOmniRequestOutput.multimodal_output["metadata"]["ar_diffusion"]. - The consumer parses the standard envelope. The session commits its reducer state, prompt, controls, event IDs, and next chunk index only when returned metadata exactly equals the submitted tick snapshot.
The metadata envelope is:
{
"metadata": {
"ar_diffusion": {
"session_id": "world-7",
"request_id": "ar-world-7-3-...",
"chunk_index": 3,
"applied_event_ids": [10, 11]
}
}
}
Transaction and failure semantics¶
Accepted events are not acknowledged as applied until the model output and metadata are validated. Reducer prepare() is speculative; reducer commit() and the session snapshot form one logical commit. A model, metadata, or reducer failure leaves the chunk index and queued events uncommitted and transitions the session to FAILED.
After a tick failure, worker state is closed because a forward may have partially modified or already released KV. The same chunk is not retried in place. A caller must explicitly reset the session, starting again at chunk zero, or close it.
Explicit close follows this state machine:
ACTIVE or FAILED
|
v
CLOSING ---- worker cleanup succeeds ----> CLOSED
|
+------- worker cleanup fails --------> CLEANUP_FAILED
|
+-- close retry --> CLOSING
CLEANUP_FAILED is a local tombstone. The session manager retains it, rejects creation of the same session_id, and rejects reset. Only an idempotent close retry that succeeds removes the session. Transport disconnect uses the same path.
KV ownership and capacity¶
ARDiffusionKVCacheSpec is the pipeline-owned immutable geometry contract. It declares:
- TP-local self-attention geometry, AR block size, sink, and recent window;
- logical KV branches and their worker-local slots;
- fixed-length cross-attention caches;
- scratch tokens required by an uncommitted forward;
- persistent model-owned CUDA bytes per resident session; and
- the requested resident-session capacity.
The runner computes:
required(capacity) =
scratch
+ managed_self_attention_pages(capacity)
+ capacity * cross_attention_bytes_per_session
+ capacity * model_owned_state_bytes_per_session
gpu_memory_fraction is a soft expansion budget:
configured_budget = available_device_bytes * gpu_memory_fraction
effective_budget = max(configured_budget, required(1))
If required(1) exceeds actual available device memory, initialization fails. Otherwise the runner always admits one viable session and uses the configured fraction to select additional resident sessions up to the model-declared cap. All model-owned reservations are deducted before the paged self-attention pool is allocated.
LingBot advances a session-owned causal Wan VAE encoder for each condition block: the opening pixel frame is followed by four zero pixel frames per later latent frame. It retains one current condition block, with separate committed and in-flight encoder histories so a failed block cannot overwrite committed context. The state reservation includes both encoder histories, derived from the encoder's convolution input grids, and the streaming decoder cache. These caches are released through the same close/reset hooks. Condition encoding currently requires the unpatched, non-tiled Wan encoder path; the existing decoder's tiling fallback remains independent. Real-weight GPU parity with a full-horizon encode still requires validation because the pointwise posterior projection now runs per block.
DreamZero reports the measured 603 MiB per-session Wan VAE causal-convolution state.
Concurrency and routing limits¶
The current runner executes one request at a time:
max_num_seqsmust be one;- request batching and diffusion step execution are rejected;
- one AR block is generated per engine request; and
- an AR-Diffusion stage must have exactly one replica.
The single-replica restriction is intentional. Multi-replica support requires session-affine routing so every tick, reset, and close reaches the worker that owns the session. Random or round-robin routing is not correct.
Multiple resident sessions may share one runner. Capacity selection accounts for all persistent pools, and the runner performs LRU eviction through the same model lifecycle hook used by explicit close.
Model boundary¶
The generic runtime transports controls as opaque ARDiffusionControlInput(track, schema, data) values. Model adapters validate and reduce those values. This keeps target/velocity, key state, pose, intrinsics, and other model-specific conventions out of the session protocol.
For LingBot World v2:
lingbot_world/actions.pyparses key-state and trajectory controls;lingbot_world/camera.pyloads or constructs camera trajectories and Plücker embeddings;lingbot_world/transformer.pyimplements checkpoint-compatible causal attention; andlingbot_world/pipeline.pyconstructs conditioning, owns small non-KV session state, and produces one block plus the standard metadata envelope.
Ulysses sequence parallelism¶
LingBot supports pure Ulysses sequence parallelism for both direct and AR-Diffusion execution. With Ulysses degree greater than one, hidden tokens, camera features, token-expanded timestep modulation, and RoPE tables are sharded together. Without SP, timestep modulation retains the frame-broadcast path. Self-attention performs the sequence-to-head all-to-all before reading or writing paged KV, while static text K/V uses the same local head shard. Text K/V shards own compact storage; cross-attention exchanges query/output layouts to use these shards. This retains the shared cache geometry and reduces text K/V storage at the cost of two all-to-all calls per layer.
The output head projects local tokens before gathering the flow values. For the 14B model this reduces the gathered width from 5120 to 64; frame modulation uses each shard's global token offset, including shards that split a frame. Only ulysses_mode="strict" is supported; advanced_uaa, Ring, and AllGather-KV modes remain unsupported for this model.
Non-goals¶
This contract does not currently provide:
- a public HTTP or WebSocket request schema (the structured-interaction frontend is tracked separately in #5527);
- camera/action semantics shared by every world model;
- session migration or replication across workers;
- retry of an ambiguous partially executed chunk;
- cross-stage KV transfer; or
- stateful streaming VAE decode.
Those features can be layered on top without adding LingBot-specific identity or lifecycle fields to the generic runtime.
Required regression coverage¶
CPU contract tests cover:
- separation of tick correlation IDs from internal engine routing IDs;
- event ordering, metadata equality, and reducer fail-closed commit;
- close/disconnect failure tombstones and cleanup retry;
- capacity one under the default fraction for shipped LingBot geometry;
- capacity expansion under a tuned fraction;
- DreamZero model-owned state accounting;
- runner session reuse, reset, close, LRU eviction, and forward cleanup; and
- LingBot model registration, imports, camera/action reduction, and one-block output metadata.
GPU validation must additionally exercise real model loading, action input, prompt switching, rolling-window replacement, reset, close, finite output latents, and exact metadata/request identity.