Skip to content

vllm_omni.diffusion.models.pi05.processor_pi05

Preprocessing for the π0.5 VLA model.

Converts a raw robot observation (multi-camera images + language instruction + proprioceptive state) into the tensors Pi05ForActionPrediction.sample_actions consumes.

The defining π0.5 difference: the state is not projected by a state_proj layer. It is normalized to [-1, 1], discretized into state_num_bins bins, and serialized into the language prompt::

"Task: <instruction>, State: <b0> <b1> ... <bN>;\nAction: "

so sample_actions receives no state tensor at all.

Normalization must precede discretization. The discretizer bins over [-1, 1] and assumes the state is already in that range. Reversed, every bin index is wrong and nothing raises.

Reference
  • OpenPI: openpi/src/openpi/shared/image_tools.py (resize_with_pad)
  • OpenPI: openpi/src/openpi/models/pi0_config.py (PaliGemmaTokenizer.tokenize)
  • LeRobot: lerobot/src/lerobot/policies/pi05/processor_pi05.py

NORM_EPS module-attribute

NORM_EPS = 1e-08

logger module-attribute

logger = init_logger(__name__)

NormStats dataclass

The affine map a checkpoint's statistics define.

mean_std carries mean and std; the two range modes carry the lower and upper bound (min/max, or the q01/q99 quantiles π0.5 ships).

lower instance-attribute

lower: Tensor

mode instance-attribute

mode: str

upper instance-attribute

upper: Tensor

Pi05ImageProcessor

Minimal image preprocessor: image → normalized + padded [-1,1] tensor.

image_size instance-attribute

image_size = image_size

make_empty_image

make_empty_image() -> Tensor

Fill tensor for an unused camera slot — pure -1, matches OpenPI/LeRobot.

preprocess_single

preprocess_single(image: Any) -> Tensor

Convert one LeRobot-domain image to a normalized model tensor.

PIL and uint8 inputs have domain [0, 255]. Floating ndarray/tensor inputs must be finite and in [0, 1]. Values are converted deterministically to [-1, 1]; the pixel content is never used to guess its input domain.

Pi05Processor

Everything between the OpenPI wire and the model.

Owns the normalization statistics and the relative-action transform, so the pipeline holds no preprocessing state of its own.

config instance-attribute

config = config

device instance-attribute

device = device

relative_actions instance-attribute

relative_actions = Pi05RelativeActions(
    enabled=config.use_relative_actions,
    exclude_joints=config.relative_exclude_joints,
    action_names=config.action_feature_names,
    max_action_dim=config.max_action_dim,
)

tokenizer instance-attribute

tokenizer = tokenizer

build_model_inputs

build_model_inputs(robot_obs: dict)

Raw observation → sample_actions inputs.

No state tensor comes back: π0.5 carries the state inside lang_tokens.

build_model_outputs

build_model_outputs(
    actions: Tensor, robot_obs: dict
) -> ndarray

Model output → the action chunk the robot receives.

AbsoluteActionsProcessorStep runs after unnormalization, because a relative-action checkpoint's statistics are computed in relative space. The raw state is the one the prompt encoded, before normalization.

Pi05RelativeActions

The relative/absolute action transform, as a single paired object.

LeRobot builds one RelativeActionsProcessorStep and hands the same instance to AbsoluteActionsProcessorStep; the two directions must agree on enabled and on which dimensions are excluded, so they are one object here too.

Deviation from LeRobot, on purpose. LeRobot's step keeps the reference state on self between the pre- and post-pass. That is safe for a single-threaded training loop and unsafe for a server: two in-flight requests would share one reference state and silently corrupt each other's actions. Here the state is passed explicitly to :meth:to_absolute, so the object stays immutable after construction and is safe to share across requests.

Transform (LeRobot / OpenPI): relative = action - state on the way in, absolute = relative + state on the way out, applied only to the dimensions not named in exclude_joints. Gripper open/close is an absolute command, which is why it is excluded by default.

action_names instance-attribute

action_names = list(action_names) if action_names else None

enabled instance-attribute

enabled = bool(enabled)

exclude_joints instance-attribute

exclude_joints = list(exclude_joints or [])

max_action_dim instance-attribute

max_action_dim = int(max_action_dim)

num_relative_dims property

num_relative_dims: int

relative_mask instance-attribute

relative_mask = mask

to_absolute

to_absolute(actions: Tensor, state: Any) -> Tensor

relative → absolute. Output side (step 2 of the post-pipeline).

state must be the raw state — the same one the model was given before normalization — because relative actions live in raw action space.

to_relative

to_relative(actions: Tensor, state: Any) -> Tensor

absolute → relative. Input side (step 3).

Not used on the inference path — there are no input actions to convert at serving time — but it is what the transform means, and the parity test exercises it as the inverse of :meth:to_absolute.

apply_norm

apply_norm(
    x: Tensor,
    stats: NormStats | None,
    *,
    inverse: bool = False,
) -> Tensor

Map raw units to the model's normalized space, or back with inverse.

Actions are padded to max_action_dim while the statistics cover only the real width, so the tail passes through untouched. The eps rules are LeRobot's: mean_std always divides by std + eps, and the range modes substitute eps only for an exactly zero range.

as_state_vector

as_state_vector(raw_state: Any, state_dim: int) -> ndarray

Coerce a request's raw state to (state_dim,) float32, or raise.

LeRobot's Pi05PrepareStateTokenizerProcessorStep discretizes whatever width it is handed and never pads to max_state_dim, so zero-filling a 7-dim state up to 32 would append 25 state tokens the checkpoint never saw in training.

build_norm_stats

build_norm_stats(
    norm_stats: dict | None, key: str
) -> NormStats | None

Read norm_stats[key] into tensors, or None when it carries none.

The mode comes from the entry. A LeRobot state_dict ships mean, std, min, max, q01 and q99 at once, so the statistic names present cannot select it. Missing modes default to quantile, which is what LeRobot defaults π0.5's STATE and ACTION to (π0 defaults to mean_std).

build_pi05_prompt

build_pi05_prompt(
    *,
    task: str,
    normalized_state: ndarray,
    state_num_bins: int = DEFAULT_STATE_NUM_BINS,
) -> str

Build the π0.5 prompt: instruction + serialized discretized state.

Matches LeRobot's Pi05PrepareStateTokenizerProcessorStep, including the task cleanup (strip, _ → space, newline → space), the exact template and the state values it serializes. The template already ends in a newline, so — unlike π0 — there is no separate newline-appending step.

normalized_state comes from Pi05ForActionPrediction._normalize_state.

discretize_state

discretize_state(
    state: ndarray,
    *,
    num_bins: int = DEFAULT_STATE_NUM_BINS,
) -> ndarray

Discretize a [-1, 1] state into num_bins integer bins.

Byte-for-byte LeRobot's processor_pi05.py::

np.digitize(state_np, bins=np.linspace(-1, 1, 256 + 1)[:-1]) - 1

A state below -1 lands in bin -1, and that negative bin is part of the contract: the checkpoint was trained with " -1" in the state prompt for those dimensions. Clipping it to 0 changes the tokens the model sees, so this must not clip.

bins stays float64, matching LeRobot's default linspace dtype, so boundary values fall on the same side.

Assumes the state is already normalized — see the module docstring.

pil_image_to_tensor

pil_image_to_tensor(image: Image) -> Tensor

PIL → (1, C, H, W) float32 in [-1, 1] (SigLIP normalization).

resize_with_pad

resize_with_pad(
    images: Tensor,
    target_height: int,
    target_width: int,
    mode: str = "bilinear",
) -> Tensor

Resize (B, C, H, W) images to the target shape, preserving aspect ratio with -1 padding on the short side.

Matches openpi image_tools.resize_with_pad_torch — the clamp to [-1, 1] is what lets the padded region blend with SigLIP-normalized pixels without adding signal at the boundary.

tokenize_prompt

tokenize_prompt(
    tokenizer,
    text: str,
    max_token_len: int = DEFAULT_MAX_TOKEN_LEN,
)

Return (input_ids, attention_mask) lists, length exactly max_token_len.

padding="max_length" is what makes the prefix a constant shape: the text segment is always max_token_len tokens regardless of the instruction, so only attention_mask.sum() varies per request.