Skip to content

vllm_omni.diffusion.models.pi05

π0.5 VLA model for vllm-omni.

PaliGemma (SigLIP vision + Gemma-2B LM) + Gemma-300M action expert with AdaRMS timestep conditioning + a flow-matching action head. Outputs a continuous action chunk [horizon, action_dim] rather than tokens.

Differs from π0 in five places: AdaRMS timestep conditioning (time_mlp_* instead of action_time_mlp_*), AdaRMS norms throughout the action expert, a 200-token tokenizer budget, a discretized state carried in the prompt (no state_proj), and relative-action support.

Modules:

Name Description
config

Config surface for the π0.5 VLA model in vllm-omni.

modeling_pi05

Inference-only π0.5 VLA math kernel for vllm-omni.

pipeline_pi05

π0.5 VLA pipeline for vllm-omni.

processor_pi05

Preprocessing for the π0.5 VLA model.

GemmaVariantConfig

depth instance-attribute

depth = depth

head_dim instance-attribute

head_dim = head_dim

mlp_dim instance-attribute

mlp_dim = mlp_dim

num_heads instance-attribute

num_heads = num_heads

num_kv_heads instance-attribute

num_kv_heads = num_kv_heads

width instance-attribute

width = width

PaliGemmaWithActionExpertPi05

Bases: Module

Dual-backbone transformer: PaliGemma (Gemma 2B) + AdaRMS expert (300M).

Same two-mode dispatch as π0 (prefix_only / suffix_only), with one structural change: after building a stock GemmaForCausalLM expert, every norm in it is swapped for a :class:Pi05AdaRMSNorm carrying a dense conditioning projection.

Swapping in place — rather than subclassing GemmaModel as #4419 does — keeps the module tree, and therefore the checkpoint key layout, identical to the expert's stock layout apart from the norms themselves.

adarms_cond_dim instance-attribute

adarms_cond_dim = action_expert_config.width

gemma_expert instance-attribute

gemma_expert = GemmaForCausalLM(
    config=action_expert_config_hf
)

paligemma instance-attribute

paligemma = PaliGemmaForConditionalGeneration(
    config=vlm_config_hf
)

embed_image

embed_image(pixel_values: Tensor) -> Tensor

Encode images with SigLIP vision tower + PaliGemma projector.

The two steps are run explicitly rather than via PaliGemmaModel.get_image_features because that helper divides the projector output by sqrt(text_hidden_size). Being explicit keeps the scale unambiguous and matches π0 exactly (SigLIP is unchanged in π0.5).

embed_language_tokens

embed_language_tokens(tokens: Tensor) -> Tensor

Embed language tokens, returning the * sqrt(hidden)-scaled embedding.

The scaling location moved across transformers releases: at ≤5.3 the normalizer lives inside GemmaModel.forward (which we bypass), and at ≥5.4 GemmaTextScaledWordEmbedding self-applies it. Detect and avoid double-scaling.

forward

forward(
    attention_mask: Tensor | None = None,
    position_ids: LongTensor | None = None,
    past_key_values=None,
    inputs_embeds: list[Tensor | None] | None = None,
    use_cache: bool = False,
    adarms_cond: Tensor | None = None,
)

Dispatch to prefix_only / suffix_only and return ([prefix_out, suffix_out], past_key_values_or_None).

Pi05AdaRMSNorm

Bases: Module

Adaptive RMSNorm conditioned on the flow-matching timestep.

π0 conditions on time by concatenating a time embedding onto each action embedding. π0.5 instead feeds the time embedding into every action-expert norm, which produces a per-layer (scale, shift, gate) triple::

y    = norm(x) * (1 + scale) + shift
out  = residual + gate * sublayer(y)

dense is zero-initialized, so an untrained model starts as the identity modulation with a closed gate — matching OpenPI's parameterization.

Note the shape of the unconditioned branch: normed * (1 + weight) with weight zero-initialized, which is exactly transformers' GemmaRMSNorm. That equivalence is why only the expert norms need replacing here and the PaliGemma prefix can keep stock Gemma layers.

cond_dim instance-attribute

cond_dim = cond_dim

dense instance-attribute

dense = nn.Linear(cond_dim, dim * 3, bias=True)

dim instance-attribute

dim = dim

eps instance-attribute

eps = eps

weight instance-attribute

weight = None

forward

forward(
    x: Tensor, cond: Tensor | None = None
) -> tuple[Tensor, Tensor | None]

Return (normed, gate); gate is None in the unconditioned case.

Pi05Config dataclass

π0.5 VLA config (dataclass, not an HF PretrainedConfig).

action_dim class-attribute instance-attribute

action_dim: int = field(init=False)

action_expert_variant class-attribute instance-attribute

action_expert_variant: str = 'gemma_300m'

action_feature_names class-attribute instance-attribute

action_feature_names: list[str] | None = None

chunk_size class-attribute instance-attribute

chunk_size: int = 50

dtype class-attribute instance-attribute

dtype: str = 'float32'

image_feature_keys class-attribute instance-attribute

image_feature_keys: list[str] | None = None

image_key_map class-attribute instance-attribute

image_key_map: dict[str, str] = field(default_factory=dict)

image_resolution class-attribute instance-attribute

image_resolution: tuple[int, int] = (224, 224)

input_features class-attribute instance-attribute

input_features: dict[str, Any] = field(default_factory=dict)

max_action_dim class-attribute instance-attribute

max_action_dim: int = 32

max_cameras class-attribute instance-attribute

max_cameras: int = 3

max_period class-attribute instance-attribute

max_period: float = 4.0

max_state_dim class-attribute instance-attribute

max_state_dim: int = 32

min_period class-attribute instance-attribute

min_period: float = 0.004

n_action_steps class-attribute instance-attribute

n_action_steps: int = 50

norm_stats class-attribute instance-attribute

norm_stats: dict | None = None

num_inference_steps class-attribute instance-attribute

num_inference_steps: int = 10

output_features class-attribute instance-attribute

output_features: dict[str, Any] = field(
    default_factory=dict
)

paligemma_variant class-attribute instance-attribute

paligemma_variant: str = 'gemma_2b'

policy_server_config class-attribute instance-attribute

policy_server_config: dict[str, Any] = field(
    default_factory=dict
)

relative_exclude_joints class-attribute instance-attribute

relative_exclude_joints: list[str] = field(
    default_factory=lambda: ["gripper"]
)

state_dim class-attribute instance-attribute

state_dim: int = field(init=False)

state_num_bins class-attribute instance-attribute

state_num_bins: int = 256

tokenizer_max_length class-attribute instance-attribute

tokenizer_max_length: int = 200

use_relative_actions class-attribute instance-attribute

use_relative_actions: bool = False

from_model_config classmethod

from_model_config(
    model_config: dict[str, Any] | None,
) -> Pi05Config

Build from a config dict (LeRobot config.json or deploy yaml).

from_pretrained classmethod

from_pretrained(checkpoint_dir: str | Path) -> Pi05Config

Build from a checkpoint directory's config.json.

Normalization stats are not in config.json. LeRobot keeps them in the processor sidecar, so they are loaded separately and backfilled here — see :func:load_lerobot_norm_stats.

Pi05ForActionPrediction

Bases: Module

π0.5 VLA model for robot action prediction via flow matching.

Inference flow
  1. Embed prefix (images + language, where the language already carries the discretized state) → prefix tokens.
  2. Forward prefix through PaliGemma → layer-wise KV cache.
  3. For each denoising step t = 1.0, 1-dt, ..., 0: a. Embed the timestep → an AdaRMS conditioning vector. b. Embed the suffix (action tokens only). c. Forward the suffix through the AdaRMS action expert. d. x_t = x_t + dt * v_t (Euler integration).
  4. Return x_0 as the predicted action chunk.

action_dim instance-attribute

action_dim = getattr(
    config, "max_action_dim", DEFAULT_ACTION_DIM
)

action_horizon instance-attribute

action_horizon = getattr(
    config, "chunk_size", DEFAULT_ACTION_HORIZON
)

action_in_proj instance-attribute

action_in_proj = nn.Linear(
    self.action_dim, self.expert_width
)

action_out_proj instance-attribute

action_out_proj = nn.Linear(
    self.expert_width, self.action_dim
)

config instance-attribute

config = config

expert_width instance-attribute

expert_width = expert_config.width

max_state_dim instance-attribute

max_state_dim = getattr(
    config, "max_state_dim", self.action_dim
)

num_inference_steps instance-attribute

num_inference_steps = getattr(
    config,
    "num_inference_steps",
    DEFAULT_NUM_INFERENCE_STEPS,
)

paligemma_with_expert instance-attribute

paligemma_with_expert = PaliGemmaWithActionExpertPi05(
    vlm_config, expert_config
)

time_mlp_in instance-attribute

time_mlp_in = nn.Linear(
    self.expert_width, self.expert_width
)

time_mlp_out instance-attribute

time_mlp_out = nn.Linear(
    self.expert_width, self.expert_width
)

vlm_width instance-attribute

vlm_width = vlm_config.width

denoise_step

denoise_step(
    prefix_pad_masks: Tensor,
    past_key_values,
    x_t: Tensor,
    timestep: Tensor,
) -> Tensor

Apply one flow-matching denoising step: predict v_t from x_t.

Signature differs from π0's by exactly one argument: no state.

embed_prefix

embed_prefix(
    images: list[Tensor],
    image_masks: list[Tensor],
    lang_tokens: Tensor,
    lang_masks: Tensor,
) -> tuple[Tensor, Tensor, Tensor]

Build prefix embeddings, per-token padding mask, and AR mask.

Prefix tokens form [img_cam_0..., ..., lang_tokens...] with fully bidirectional attention (all-zero att_masks). Identical to π0 — the state is inside lang_tokens, so nothing here changes shape-wise.

Cameras are embedded one at a time; each call is (B, 3, 224, 224). The number of slots is fixed by config.max_cameras for the deployed model; missing cameras occupy their slot with a false image mask.

embed_suffix

embed_suffix(
    noisy_actions: Tensor, timestep: Tensor
) -> tuple[Tensor, Tensor, Tensor, Tensor]

Build the suffix: action tokens only, plus the AdaRMS condition.

π0's suffix is [state_token, action_tokens×H] with an AR mask of [1, 1, 0...]. π0.5 has no state token, so the suffix is [action_tokens×H] and the mask is [1] + [0]*(H-1): the first action token opens a causal block and the rest attend bidirectionally within it.

embed_timestep

embed_timestep(timestep: Tensor) -> Tensor

Timestep → AdaRMS conditioning vector (B, expert_width).

silu(time_mlp_out(silu(time_mlp_in(sinusoid(t))))). The trailing SiLU is part of the reference implementation — dropping it is a silent numerical error, not a crash.

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
    *,
    strict: bool = True,
)

Load and audit a LeRobot π0.5 safetensors checkpoint.

Same remap rules as π0 (strip the model. prefix, flatten→nested PaliGemma submodules, tied lm_head → embed_tokens, version-robust SigLIP nesting), plus two π0.5-specific ones:

  • action_time_mlp_{in,out} → time_mlp_{in,out}: some checkpoints were exported under the π0 parameter names.
  • state_proj.* is reported, not silently dropped. A π0.5 checkpoint should not contain it; its presence usually means a π0 checkpoint was pointed at the π0.5 model class, which would otherwise run happily with a randomly-initialized action expert.

The action-expert norms are AdaRMS here, so they expose dense.weight / dense.bias and no plain weight. A checkpoint that carries a plain expert-norm weight is a π0-shaped checkpoint; that too is rejected rather than skipped. strict=False exists only for focused remapping unit tests that intentionally provide a partial state dict; the serving path always uses the strict default.

sample_actions

sample_actions(
    images: list[Tensor],
    image_masks: list[Tensor],
    lang_tokens: Tensor,
    lang_masks: Tensor,
    noise: Tensor | None = None,
    num_steps: int | None = None,
    generator: Generator | list[Generator] | None = None,
) -> Tensor

Generate an action chunk via iterative flow-matching denoising.

Convention: t=1 is noise, t=0 is the target — opposite of the published π0 paper but matching both OpenPI and LeRobot.

Takes no state: π0.5's state rides inside lang_tokens.

Pi05ImageProcessor

Minimal image preprocessor: image → normalized + padded [-1,1] tensor.

image_size instance-attribute

image_size = image_size

make_empty_image

make_empty_image() -> Tensor

Fill tensor for an unused camera slot — pure -1, matches OpenPI/LeRobot.

preprocess_single

preprocess_single(image: Any) -> Tensor

Convert one LeRobot-domain image to a normalized model tensor.

PIL and uint8 inputs have domain [0, 255]. Floating ndarray/tensor inputs must be finite and in [0, 1]. Values are converted deterministically to [-1, 1]; the pixel content is never used to guess its input domain.

Pi05Pipeline

Bases: Module

π0.5 VLA pipeline: raw robot obs → continuous action chunk.

Registered as "Pi05Pipeline" in the diffusion registry.

config instance-attribute

config = self._build_config(od_config)

model instance-attribute

model = self._initialize_model()

model_dir instance-attribute

model_dir = self._resolve_model_dir(od_config.model)

od_config instance-attribute

od_config = od_config

prefix instance-attribute

prefix = prefix

processor instance-attribute

processor = Pi05Processor(
    self.config, self.tokenizer, self._device
)

tokenizer instance-attribute

tokenizer = self._load_tokenizer()

tokenizer_source instance-attribute

tokenizer_source = str(
    custom_args.get(
        "tokenizer", self._resolve_tokenizer_source()
    )
)

forward

forward(
    req: OmniDiffusionRequest, **kwargs
) -> DiffusionOutput

has_real_checkpoint

has_real_checkpoint() -> bool

load_weights

load_weights(weights=())

No-op for the diffusion loader: π0.5 self-loads its checkpoint.

Pi05Processor

Everything between the OpenPI wire and the model.

Owns the normalization statistics and the relative-action transform, so the pipeline holds no preprocessing state of its own.

config instance-attribute

config = config

device instance-attribute

device = device

relative_actions instance-attribute

relative_actions = Pi05RelativeActions(
    enabled=config.use_relative_actions,
    exclude_joints=config.relative_exclude_joints,
    action_names=config.action_feature_names,
    max_action_dim=config.max_action_dim,
)

tokenizer instance-attribute

tokenizer = tokenizer

build_model_inputs

build_model_inputs(robot_obs: dict)

Raw observation → sample_actions inputs.

No state tensor comes back: π0.5 carries the state inside lang_tokens.

build_model_outputs

build_model_outputs(
    actions: Tensor, robot_obs: dict
) -> ndarray

Model output → the action chunk the robot receives.

AbsoluteActionsProcessorStep runs after unnormalization, because a relative-action checkpoint's statistics are computed in relative space. The raw state is the one the prompt encoded, before normalization.

Pi05RelativeActions

The relative/absolute action transform, as a single paired object.

LeRobot builds one RelativeActionsProcessorStep and hands the same instance to AbsoluteActionsProcessorStep; the two directions must agree on enabled and on which dimensions are excluded, so they are one object here too.

Deviation from LeRobot, on purpose. LeRobot's step keeps the reference state on self between the pre- and post-pass. That is safe for a single-threaded training loop and unsafe for a server: two in-flight requests would share one reference state and silently corrupt each other's actions. Here the state is passed explicitly to :meth:to_absolute, so the object stays immutable after construction and is safe to share across requests.

Transform (LeRobot / OpenPI): relative = action - state on the way in, absolute = relative + state on the way out, applied only to the dimensions not named in exclude_joints. Gripper open/close is an absolute command, which is why it is excluded by default.

action_names instance-attribute

action_names = list(action_names) if action_names else None

enabled instance-attribute

enabled = bool(enabled)

exclude_joints instance-attribute

exclude_joints = list(exclude_joints or [])

max_action_dim instance-attribute

max_action_dim = int(max_action_dim)

num_relative_dims property

num_relative_dims: int

relative_mask instance-attribute

relative_mask = mask

to_absolute

to_absolute(actions: Tensor, state: Any) -> Tensor

relative → absolute. Output side (step 2 of the post-pipeline).

state must be the raw state — the same one the model was given before normalization — because relative actions live in raw action space.

to_relative

to_relative(actions: Tensor, state: Any) -> Tensor

absolute → relative. Input side (step 3).

Not used on the inference path — there are no input actions to convert at serving time — but it is what the transform means, and the parity test exercises it as the inverse of :meth:to_absolute.

UnsupportedCheckpointCapabilityError

Bases: ValueError

A checkpoint declares a capability this implementation does not consume.

Raised at load time rather than silently ignored: every one of these capabilities changes what a correct action chunk looks like, and none of them is visible in the weights alone.

apply_norm

apply_norm(
    x: Tensor,
    stats: NormStats | None,
    *,
    inverse: bool = False,
) -> Tensor

Map raw units to the model's normalized space, or back with inverse.

Actions are padded to max_action_dim while the statistics cover only the real width, so the tail passes through untouched. The eps rules are LeRobot's: mean_std always divides by std + eps, and the range modes substitute eps only for an exactly zero range.

build_norm_stats

build_norm_stats(
    norm_stats: dict | None, key: str
) -> NormStats | None

Read norm_stats[key] into tensors, or None when it carries none.

The mode comes from the entry. A LeRobot state_dict ships mean, std, min, max, q01 and q99 at once, so the statistic names present cannot select it. Missing modes default to quantile, which is what LeRobot defaults π0.5's STATE and ACTION to (π0 defaults to mean_std).

build_pi05_prompt

build_pi05_prompt(
    *,
    task: str,
    normalized_state: ndarray,
    state_num_bins: int = DEFAULT_STATE_NUM_BINS,
) -> str

Build the π0.5 prompt: instruction + serialized discretized state.

Matches LeRobot's Pi05PrepareStateTokenizerProcessorStep, including the task cleanup (strip, _ → space, newline → space), the exact template and the state values it serializes. The template already ends in a newline, so — unlike π0 — there is no separate newline-appending step.

normalized_state comes from Pi05ForActionPrediction._normalize_state.

create_sinusoidal_pos_embedding

create_sinusoidal_pos_embedding(
    time: Tensor,
    dimension: int,
    min_period: float = 0.004,
    max_period: float = 4.0,
    device: device = None,
) -> Tensor

Compute a sine/cosine positional embedding for scalar timesteps.

Ref: openpi/models_pytorch/pi0_pytorch.py create_sinusoidal_pos_embedding

discretize_state

discretize_state(
    state: ndarray,
    *,
    num_bins: int = DEFAULT_STATE_NUM_BINS,
) -> ndarray

Discretize a [-1, 1] state into num_bins integer bins.

Byte-for-byte LeRobot's processor_pi05.py::

np.digitize(state_np, bins=np.linspace(-1, 1, 256 + 1)[:-1]) - 1

A state below -1 lands in bin -1, and that negative bin is part of the contract: the checkpoint was trained with " -1" in the state prompt for those dimensions. Clipping it to 0 changes the tokens the model sees, so this must not clip.

bins stays float64, matching LeRobot's default linspace dtype, so boundary values fall on the same side.

Assumes the state is already normalized — see the module docstring.

get_gemma_config

get_gemma_config(variant: str) -> GemmaVariantConfig

get_pi05_post_process_func

get_pi05_post_process_func(od_config: OmniDiffusionConfig)

π0.5 returns actions directly; post-processing is identity.

make_att_2d_masks

make_att_2d_masks(
    pad_masks: Tensor, att_masks: Tensor
) -> Tensor

Build a 2D attention mask from a padding mask and an autoregressive mask.

Ref: openpi/models_pytorch/pi0_pytorch.py make_att_2d_masks

pil_image_to_tensor

pil_image_to_tensor(image: Image) -> Tensor

PIL → (1, C, H, W) float32 in [-1, 1] (SigLIP normalization).

prepare_attention_masks_4d

prepare_attention_masks_4d(att_2d_masks: Tensor) -> Tensor

Convert (B, S, S) bool masks to (B, 1, S, S) float masks.

True → 0.0 (attend), False → OPENPI_ATTENTION_MASK_VALUE.

resize_with_pad

resize_with_pad(
    images: Tensor,
    target_height: int,
    target_width: int,
    mode: str = "bilinear",
) -> Tensor

Resize (B, C, H, W) images to the target shape, preserving aspect ratio with -1 padding on the short side.

Matches openpi image_tools.resize_with_pad_torch — the clamp to [-1, 1] is what lets the padded region blend with SigLIP-normalized pixels without adding signal at the boundary.

resolve_excluded_action_indices

resolve_excluded_action_indices(
    exclude_joints: list[str] | None,
    action_names: list[str] | None,
) -> list[int]

Map relative_exclude_joints names onto action-vector indices.

Matching is exact name first, then substring (a checkpoint may name the gripper dimension gripper_position while the config just says gripper). Single source of truth for both the config-time validation and the runtime mask in processor_pi05.Pi05RelativeActions.

Returns an empty list when there is nothing to exclude. Raises when a name cannot be resolved — an unresolvable exclusion would otherwise silently become "make this dimension relative too".

tokenize_prompt

tokenize_prompt(
    tokenizer,
    text: str,
    max_token_len: int = DEFAULT_MAX_TOKEN_LEN,
)

Return (input_ids, attention_mask) lists, length exactly max_token_len.

padding="max_length" is what makes the prefix a constant shape: the text segment is always max_token_len tokens regardless of the instruction, so only attention_mask.sum() varies per request.