Skip to content

vllm_omni.diffusion.models.pi0.processor_pi0

Preprocessing for the π0 VLA model.

Converts a raw robot observation (multi-camera images + language instruction + proprioceptive state) into the tensors that Pi0ForActionPrediction.sample_actions consumes:

  • images : list of (1, 3, 224, 224) float tensors in [-1, 1], one per camera slot (real cameras first, then -1-filled padding) in the checkpoint's image_feature_keys order
  • image_masks : list of (1,) bool tensors, True = real camera
  • lang_tokens : (1, L) long, PaliGemma tokenizer, \n-terminated
  • lang_masks : (1, L) bool
  • state : (1, max_state_dim) float32, zero-padded

Image/prompt preprocessing matches OpenPI and LeRobot bit-for-bit — see the parity test (test_pi0_parity.py). A set of stateless helpers the diffusion pipeline calls directly (the DreamZero contract: the pipeline owns its preprocessing).

Reference
  • OpenPI: openpi/src/openpi/shared/image_tools.py (resize_with_pad)
  • OpenPI: openpi/src/openpi/models_pytorch/preprocessing_pytorch.py
  • LeRobot: lerobot/src/lerobot/policies/pi0/modeling_pi0.py (Pi0NewLineProcessor)

PI0_IMAGE_SIZE module-attribute

PI0_IMAGE_SIZE = 224

PI0_IMAGE_TOKEN_INDEX module-attribute

PI0_IMAGE_TOKEN_INDEX = 257152

PI0_MAX_CAMERAS module-attribute

PI0_MAX_CAMERAS = 3

PI0_MAX_TOKEN_LEN module-attribute

PI0_MAX_TOKEN_LEN = 48

PI0_NUM_IMAGE_TOKENS module-attribute

PI0_NUM_IMAGE_TOKENS = 256

logger module-attribute

logger = logging.getLogger(__name__)

Pi0ImageProcessor

Minimal image preprocessor: image → normalized + padded [-1,1] tensor.

image_size instance-attribute

image_size = image_size

make_empty_image

make_empty_image() -> Tensor

Fill tensor for an unused camera slot — pure -1, matches OpenPI/LeRobot.

preprocess_single

preprocess_single(image: Any) -> Tensor

Accept a PIL image, an HWC uint8/float ndarray, or a CHW tensor and return a (1, 3, image_size, image_size) float tensor in [-1, 1].

build_model_inputs

build_model_inputs(
    robot_obs: dict, config, tokenizer, device: device
)

Convert a raw robot observation into sample_actions inputs.

Reproduces the LeRobot camera ordering exactly: iterate config.image_feature_keys (the ordered camera identities from the checkpoint's input_features); for each, use the supplied image (mask True) or a -1-filled empty image (mask False). Then tokenize the prompt and zero-pad/truncate the state to max_state_dim.

Returns (images, image_masks, lang_tokens, lang_masks, state) with a leading batch dim of 1 on the tensors.

pil_image_to_tensor

pil_image_to_tensor(image: Image) -> Tensor

PIL → (1, C, H, W) float32 in [-1, 1] (SigLIP normalization).

resize_with_pad

resize_with_pad(
    images: Tensor,
    target_height: int,
    target_width: int,
    mode: str = "bilinear",
) -> Tensor

Resize (B, C, H, W) images to the target shape, preserving aspect ratio with -1 padding on the short side.

Matches openpi image_tools.resize_with_pad_torch — the clamp to [-1, 1] is what lets the padded region blend with SigLIP-normalized pixels without adding signal at the boundary.

tokenize_prompt

tokenize_prompt(
    tokenizer,
    text: str,
    max_token_len: int = PI0_MAX_TOKEN_LEN,
)

Return (input_ids, attention_mask) lists, length max_token_len.

Matches LeRobot's Pi0NewLineProcessor: append \n if the caller didn't, then run the PaliGemma tokenizer with right-padding.