vllm_omni.diffusion.models.pi0.processor_pi0 ¶
Preprocessing for the π0 VLA model.
Converts a raw robot observation (multi-camera images + language instruction + proprioceptive state) into the tensors that Pi0ForActionPrediction.sample_actions consumes:
- images : list of
(1, 3, 224, 224)float tensors in[-1, 1], one per camera slot (real cameras first, then-1-filled padding) in the checkpoint'simage_feature_keysorder - image_masks : list of
(1,)bool tensors,True= real camera - lang_tokens :
(1, L)long, PaliGemma tokenizer,\n-terminated - lang_masks :
(1, L)bool - state :
(1, max_state_dim)float32, zero-padded
Image/prompt preprocessing matches OpenPI and LeRobot bit-for-bit — see the parity test (test_pi0_parity.py). A set of stateless helpers the diffusion pipeline calls directly (the DreamZero contract: the pipeline owns its preprocessing).
Reference
- OpenPI: openpi/src/openpi/shared/image_tools.py (resize_with_pad)
- OpenPI: openpi/src/openpi/models_pytorch/preprocessing_pytorch.py
- LeRobot: lerobot/src/lerobot/policies/pi0/modeling_pi0.py (Pi0NewLineProcessor)
Pi0ImageProcessor ¶
Minimal image preprocessor: image → normalized + padded [-1,1] tensor.
make_empty_image ¶
Fill tensor for an unused camera slot — pure -1, matches OpenPI/LeRobot.
build_model_inputs ¶
build_model_inputs(
robot_obs: dict, config, tokenizer, device: device
)
Convert a raw robot observation into sample_actions inputs.
Reproduces the LeRobot camera ordering exactly: iterate config.image_feature_keys (the ordered camera identities from the checkpoint's input_features); for each, use the supplied image (mask True) or a -1-filled empty image (mask False). Then tokenize the prompt and zero-pad/truncate the state to max_state_dim.
Returns (images, image_masks, lang_tokens, lang_masks, state) with a leading batch dim of 1 on the tensors.
pil_image_to_tensor ¶
PIL → (1, C, H, W) float32 in [-1, 1] (SigLIP normalization).
resize_with_pad ¶
resize_with_pad(
images: Tensor,
target_height: int,
target_width: int,
mode: str = "bilinear",
) -> Tensor
Resize (B, C, H, W) images to the target shape, preserving aspect ratio with -1 padding on the short side.
Matches openpi image_tools.resize_with_pad_torch — the clamp to [-1, 1] is what lets the padded region blend with SigLIP-normalized pixels without adding signal at the boundary.
tokenize_prompt ¶
tokenize_prompt(
tokenizer,
text: str,
max_token_len: int = PI0_MAX_TOKEN_LEN,
)
Return (input_ids, attention_mask) lists, length max_token_len.
Matches LeRobot's Pi0NewLineProcessor: append \n if the caller didn't, then run the PaliGemma tokenizer with right-padding.