Skip to content

vllm_omni.diffusion.models.lingbot_world.actions

Realtime keyboard controls for LingBot World 2.0.

LINGBOT_CAMERA_ACTION_SCHEMA module-attribute

LINGBOT_CAMERA_ACTION_SCHEMA = 'lingbot.camera_actions.v1'

LINGBOT_CAMERA_TRAJECTORY_SCHEMA module-attribute

LINGBOT_CAMERA_TRAJECTORY_SCHEMA = (
    "lingbot.camera_trajectory.v1"
)

LINGBOT_CONTROLLER_TRANSLATION_UNIT module-attribute

LINGBOT_CONTROLLER_TRANSLATION_UNIT = 0.05

LingBotCameraActionFrames module-attribute

LingBotCameraActionFrames: TypeAlias = tuple[
    tuple[str, ...], ...
]

LingBotCameraActionScript module-attribute

LingBotCameraActionScript: TypeAlias = tuple[
    LingBotCameraActionFrames, ...
]

LingBotCameraControlReducer

Sample live state/script controls into one LingBot AR block.

State mode preserves held keys across chunks and guarantees a one-frame pulse when a press and release both arrive before the next chunk. Script mode is a finite FIFO padded with neutral frames after exhaustion.

commit

commit(prepared: ARDiffusionPreparedControls) -> None

prepare

prepare(
    *,
    current_controls: Mapping[str, ARDiffusionControlInput],
    events: Sequence[ARDiffusionSessionEvent],
    chunk_index: int,
) -> ARDiffusionPreparedControls

reset

reset() -> None

as_camera_action_frames

as_camera_action_frames(
    value: Iterable[Iterable[str]],
) -> LingBotCameraActionFrames

Restore the tuple form of one chunk's per-latent-frame key states.

as_camera_action_script

as_camera_action_script(
    value: Iterable[Iterable[Iterable[str]]],
) -> LingBotCameraActionScript

Restore the tuple form of a whole request's script.

sampling_params.extra_args round-trips through JSON, so an already validated script comes back as lists; this rebuilds the hashable tuple form without re-running validation.

camera_trajectory_from_absolute_pose

camera_trajectory_from_absolute_pose(
    poses: Tensor, *, width: int, height: int
) -> CameraTrajectory

integrate_lingbot_camera_actions

integrate_lingbot_camera_actions(
    frames: Sequence[Sequence[str]],
    *,
    width: int,
    height: int,
    initial_pose: Tensor | None = None,
    initial_pitch: float = 0.0,
) -> tuple[CameraTrajectory, float]

Convert latent-frame WASD/IJKL controls to cumulative C2W poses.

Calibration and motion constants intentionally match SGLang's LingBot adapter: movement 0.05 (LINGBOT_CONTROLLER_TRANSLATION_UNIT), pitch 4 degrees, yaw 6 degrees, pitch clamp 85 degrees, and target-resolution focal lengths of 500 pixels.

parse_lingbot_camera_action_frames

parse_lingbot_camera_action_frames(
    data: Mapping[str, Any], *, expected_frames: int
) -> LingBotCameraActionFrames

Validate the chunk-sized payload emitted by the session reducer.

parse_lingbot_camera_action_script

parse_lingbot_camera_action_script(
    script: object, *, frames_per_chunk: int
) -> LingBotCameraActionScript

Validate a request-scoped list of per-chunk camera action frames.