vllm_omni.diffusion.models.lingbot_world.actions ¶
Realtime keyboard controls for LingBot World 2.0.
LINGBOT_CAMERA_ACTION_SCHEMA module-attribute ¶
LINGBOT_CAMERA_TRAJECTORY_SCHEMA module-attribute ¶
LingBotCameraActionFrames module-attribute ¶
LingBotCameraActionScript module-attribute ¶
LingBotCameraActionScript: TypeAlias = tuple[
LingBotCameraActionFrames, ...
]
LingBotCameraControlReducer ¶
Sample live state/script controls into one LingBot AR block.
State mode preserves held keys across chunks and guarantees a one-frame pulse when a press and release both arrive before the next chunk. Script mode is a finite FIFO padded with neutral frames after exhaustion.
as_camera_action_frames ¶
as_camera_action_frames(
value: Iterable[Iterable[str]],
) -> LingBotCameraActionFrames
Restore the tuple form of one chunk's per-latent-frame key states.
as_camera_action_script ¶
as_camera_action_script(
value: Iterable[Iterable[Iterable[str]]],
) -> LingBotCameraActionScript
Restore the tuple form of a whole request's script.
sampling_params.extra_args round-trips through JSON, so an already validated script comes back as lists; this rebuilds the hashable tuple form without re-running validation.
camera_trajectory_from_absolute_pose ¶
camera_trajectory_from_absolute_pose(
poses: Tensor, *, width: int, height: int
) -> CameraTrajectory
integrate_lingbot_camera_actions ¶
integrate_lingbot_camera_actions(
frames: Sequence[Sequence[str]],
*,
width: int,
height: int,
initial_pose: Tensor | None = None,
initial_pitch: float = 0.0,
) -> tuple[CameraTrajectory, float]
Convert latent-frame WASD/IJKL controls to cumulative C2W poses.
Calibration and motion constants intentionally match SGLang's LingBot adapter: movement 0.05 (LINGBOT_CONTROLLER_TRANSLATION_UNIT), pitch 4 degrees, yaw 6 degrees, pitch clamp 85 degrees, and target-resolution focal lengths of 500 pixels.
parse_lingbot_camera_action_frames ¶
parse_lingbot_camera_action_frames(
data: Mapping[str, Any], *, expected_frames: int
) -> LingBotCameraActionFrames
Validate the chunk-sized payload emitted by the session reducer.
parse_lingbot_camera_action_script ¶
parse_lingbot_camera_action_script(
script: object, *, frames_per_chunk: int
) -> LingBotCameraActionScript
Validate a request-scoped list of per-chunk camera action frames.