vllm_omni.diffusion.models.pi05 ¶
π0.5 VLA model for vllm-omni.
PaliGemma (SigLIP vision + Gemma-2B LM) + Gemma-300M action expert with AdaRMS timestep conditioning + a flow-matching action head. Outputs a continuous action chunk [horizon, action_dim] rather than tokens.
Differs from π0 in five places: AdaRMS timestep conditioning (time_mlp_* instead of action_time_mlp_*), AdaRMS norms throughout the action expert, a 200-token tokenizer budget, a discretized state carried in the prompt (no state_proj), and relative-action support.
Modules:
| Name | Description |
|---|---|
config | Config surface for the π0.5 VLA model in vllm-omni. |
modeling_pi05 | Inference-only π0.5 VLA math kernel for vllm-omni. |
pipeline_pi05 | π0.5 VLA pipeline for vllm-omni. |
processor_pi05 | Preprocessing for the π0.5 VLA model. |
GemmaVariantConfig ¶
PaliGemmaWithActionExpertPi05 ¶
Bases: Module
Dual-backbone transformer: PaliGemma (Gemma 2B) + AdaRMS expert (300M).
Same two-mode dispatch as π0 (prefix_only / suffix_only), with one structural change: after building a stock GemmaForCausalLM expert, every norm in it is swapped for a :class:Pi05AdaRMSNorm carrying a dense conditioning projection.
Swapping in place — rather than subclassing GemmaModel as #4419 does — keeps the module tree, and therefore the checkpoint key layout, identical to the expert's stock layout apart from the norms themselves.
paligemma instance-attribute ¶
embed_image ¶
Encode images with SigLIP vision tower + PaliGemma projector.
The two steps are run explicitly rather than via PaliGemmaModel.get_image_features because that helper divides the projector output by sqrt(text_hidden_size). Being explicit keeps the scale unambiguous and matches π0 exactly (SigLIP is unchanged in π0.5).
embed_language_tokens ¶
Embed language tokens, returning the * sqrt(hidden)-scaled embedding.
The scaling location moved across transformers releases: at ≤5.3 the normalizer lives inside GemmaModel.forward (which we bypass), and at ≥5.4 GemmaTextScaledWordEmbedding self-applies it. Detect and avoid double-scaling.
forward ¶
forward(
attention_mask: Tensor | None = None,
position_ids: LongTensor | None = None,
past_key_values=None,
inputs_embeds: list[Tensor | None] | None = None,
use_cache: bool = False,
adarms_cond: Tensor | None = None,
)
Dispatch to prefix_only / suffix_only and return ([prefix_out, suffix_out], past_key_values_or_None).
Pi05AdaRMSNorm ¶
Bases: Module
Adaptive RMSNorm conditioned on the flow-matching timestep.
π0 conditions on time by concatenating a time embedding onto each action embedding. π0.5 instead feeds the time embedding into every action-expert norm, which produces a per-layer (scale, shift, gate) triple::
y = norm(x) * (1 + scale) + shift
out = residual + gate * sublayer(y)
dense is zero-initialized, so an untrained model starts as the identity modulation with a closed gate — matching OpenPI's parameterization.
Note the shape of the unconditioned branch: normed * (1 + weight) with weight zero-initialized, which is exactly transformers' GemmaRMSNorm. That equivalence is why only the expert norms need replacing here and the PaliGemma prefix can keep stock Gemma layers.
Pi05Config dataclass ¶
π0.5 VLA config (dataclass, not an HF PretrainedConfig).
action_expert_variant class-attribute instance-attribute ¶
action_expert_variant: str = 'gemma_300m'
action_feature_names class-attribute instance-attribute ¶
image_key_map class-attribute instance-attribute ¶
image_resolution class-attribute instance-attribute ¶
input_features class-attribute instance-attribute ¶
output_features class-attribute instance-attribute ¶
policy_server_config class-attribute instance-attribute ¶
relative_exclude_joints class-attribute instance-attribute ¶
from_model_config classmethod ¶
from_model_config(
model_config: dict[str, Any] | None,
) -> Pi05Config
Build from a config dict (LeRobot config.json or deploy yaml).
from_pretrained classmethod ¶
from_pretrained(checkpoint_dir: str | Path) -> Pi05Config
Build from a checkpoint directory's config.json.
Normalization stats are not in config.json. LeRobot keeps them in the processor sidecar, so they are loaded separately and backfilled here — see :func:load_lerobot_norm_stats.
Pi05ForActionPrediction ¶
Bases: Module
π0.5 VLA model for robot action prediction via flow matching.
Inference flow
- Embed prefix (images + language, where the language already carries the discretized state) → prefix tokens.
- Forward prefix through PaliGemma → layer-wise KV cache.
- For each denoising step
t = 1.0, 1-dt, ..., 0: a. Embed the timestep → an AdaRMS conditioning vector. b. Embed the suffix (action tokens only). c. Forward the suffix through the AdaRMS action expert. d.x_t = x_t + dt * v_t(Euler integration). - Return
x_0as the predicted action chunk.
action_dim instance-attribute ¶
action_dim = getattr(
config, "max_action_dim", DEFAULT_ACTION_DIM
)
action_horizon instance-attribute ¶
action_horizon = getattr(
config, "chunk_size", DEFAULT_ACTION_HORIZON
)
action_in_proj instance-attribute ¶
action_out_proj instance-attribute ¶
max_state_dim instance-attribute ¶
max_state_dim = getattr(
config, "max_state_dim", self.action_dim
)
num_inference_steps instance-attribute ¶
num_inference_steps = getattr(
config,
"num_inference_steps",
DEFAULT_NUM_INFERENCE_STEPS,
)
paligemma_with_expert instance-attribute ¶
paligemma_with_expert = PaliGemmaWithActionExpertPi05(
vlm_config, expert_config
)
denoise_step ¶
Apply one flow-matching denoising step: predict v_t from x_t.
Signature differs from π0's by exactly one argument: no state.
embed_prefix ¶
embed_prefix(
images: list[Tensor],
image_masks: list[Tensor],
lang_tokens: Tensor,
lang_masks: Tensor,
) -> tuple[Tensor, Tensor, Tensor]
Build prefix embeddings, per-token padding mask, and AR mask.
Prefix tokens form [img_cam_0..., ..., lang_tokens...] with fully bidirectional attention (all-zero att_masks). Identical to π0 — the state is inside lang_tokens, so nothing here changes shape-wise.
Cameras are embedded one at a time; each call is (B, 3, 224, 224). The number of slots is fixed by config.max_cameras for the deployed model; missing cameras occupy their slot with a false image mask.
embed_suffix ¶
embed_suffix(
noisy_actions: Tensor, timestep: Tensor
) -> tuple[Tensor, Tensor, Tensor, Tensor]
Build the suffix: action tokens only, plus the AdaRMS condition.
π0's suffix is [state_token, action_tokens×H] with an AR mask of [1, 1, 0...]. π0.5 has no state token, so the suffix is [action_tokens×H] and the mask is [1] + [0]*(H-1): the first action token opens a causal block and the rest attend bidirectionally within it.
embed_timestep ¶
Timestep → AdaRMS conditioning vector (B, expert_width).
silu(time_mlp_out(silu(time_mlp_in(sinusoid(t))))). The trailing SiLU is part of the reference implementation — dropping it is a silent numerical error, not a crash.
load_weights ¶
Load and audit a LeRobot π0.5 safetensors checkpoint.
Same remap rules as π0 (strip the model. prefix, flatten→nested PaliGemma submodules, tied lm_head → embed_tokens, version-robust SigLIP nesting), plus two π0.5-specific ones:
action_time_mlp_{in,out}→time_mlp_{in,out}: some checkpoints were exported under the π0 parameter names.state_proj.*is reported, not silently dropped. A π0.5 checkpoint should not contain it; its presence usually means a π0 checkpoint was pointed at the π0.5 model class, which would otherwise run happily with a randomly-initialized action expert.
The action-expert norms are AdaRMS here, so they expose dense.weight / dense.bias and no plain weight. A checkpoint that carries a plain expert-norm weight is a π0-shaped checkpoint; that too is rejected rather than skipped. strict=False exists only for focused remapping unit tests that intentionally provide a partial state dict; the serving path always uses the strict default.
sample_actions ¶
sample_actions(
images: list[Tensor],
image_masks: list[Tensor],
lang_tokens: Tensor,
lang_masks: Tensor,
noise: Tensor | None = None,
num_steps: int | None = None,
generator: Generator | list[Generator] | None = None,
) -> Tensor
Generate an action chunk via iterative flow-matching denoising.
Convention: t=1 is noise, t=0 is the target — opposite of the published π0 paper but matching both OpenPI and LeRobot.
Takes no state: π0.5's state rides inside lang_tokens.
Pi05ImageProcessor ¶
Minimal image preprocessor: image → normalized + padded [-1,1] tensor.
make_empty_image ¶
Fill tensor for an unused camera slot — pure -1, matches OpenPI/LeRobot.
preprocess_single ¶
preprocess_single(image: Any) -> Tensor
Convert one LeRobot-domain image to a normalized model tensor.
PIL and uint8 inputs have domain [0, 255]. Floating ndarray/tensor inputs must be finite and in [0, 1]. Values are converted deterministically to [-1, 1]; the pixel content is never used to guess its input domain.
Pi05Pipeline ¶
Bases: Module
π0.5 VLA pipeline: raw robot obs → continuous action chunk.
Registered as "Pi05Pipeline" in the diffusion registry.
processor instance-attribute ¶
processor = Pi05Processor(
self.config, self.tokenizer, self._device
)
tokenizer_source instance-attribute ¶
tokenizer_source = str(
custom_args.get(
"tokenizer", self._resolve_tokenizer_source()
)
)
load_weights ¶
No-op for the diffusion loader: π0.5 self-loads its checkpoint.
Pi05Processor ¶
Everything between the OpenPI wire and the model.
Owns the normalization statistics and the relative-action transform, so the pipeline holds no preprocessing state of its own.
relative_actions instance-attribute ¶
relative_actions = Pi05RelativeActions(
enabled=config.use_relative_actions,
exclude_joints=config.relative_exclude_joints,
action_names=config.action_feature_names,
max_action_dim=config.max_action_dim,
)
build_model_inputs ¶
build_model_inputs(robot_obs: dict)
Raw observation → sample_actions inputs.
No state tensor comes back: π0.5 carries the state inside lang_tokens.
build_model_outputs ¶
Model output → the action chunk the robot receives.
AbsoluteActionsProcessorStep runs after unnormalization, because a relative-action checkpoint's statistics are computed in relative space. The raw state is the one the prompt encoded, before normalization.
Pi05RelativeActions ¶
The relative/absolute action transform, as a single paired object.
LeRobot builds one RelativeActionsProcessorStep and hands the same instance to AbsoluteActionsProcessorStep; the two directions must agree on enabled and on which dimensions are excluded, so they are one object here too.
Deviation from LeRobot, on purpose. LeRobot's step keeps the reference state on self between the pre- and post-pass. That is safe for a single-threaded training loop and unsafe for a server: two in-flight requests would share one reference state and silently corrupt each other's actions. Here the state is passed explicitly to :meth:to_absolute, so the object stays immutable after construction and is safe to share across requests.
Transform (LeRobot / OpenPI): relative = action - state on the way in, absolute = relative + state on the way out, applied only to the dimensions not named in exclude_joints. Gripper open/close is an absolute command, which is why it is excluded by default.
to_absolute ¶
to_absolute(actions: Tensor, state: Any) -> Tensor
relative → absolute. Output side (step 2 of the post-pipeline).
state must be the raw state — the same one the model was given before normalization — because relative actions live in raw action space.
to_relative ¶
to_relative(actions: Tensor, state: Any) -> Tensor
absolute → relative. Input side (step 3).
Not used on the inference path — there are no input actions to convert at serving time — but it is what the transform means, and the parity test exercises it as the inverse of :meth:to_absolute.
UnsupportedCheckpointCapabilityError ¶
Bases: ValueError
A checkpoint declares a capability this implementation does not consume.
Raised at load time rather than silently ignored: every one of these capabilities changes what a correct action chunk looks like, and none of them is visible in the weights alone.
apply_norm ¶
Map raw units to the model's normalized space, or back with inverse.
Actions are padded to max_action_dim while the statistics cover only the real width, so the tail passes through untouched. The eps rules are LeRobot's: mean_std always divides by std + eps, and the range modes substitute eps only for an exactly zero range.
build_norm_stats ¶
Read norm_stats[key] into tensors, or None when it carries none.
The mode comes from the entry. A LeRobot state_dict ships mean, std, min, max, q01 and q99 at once, so the statistic names present cannot select it. Missing modes default to quantile, which is what LeRobot defaults π0.5's STATE and ACTION to (π0 defaults to mean_std).
build_pi05_prompt ¶
build_pi05_prompt(
*,
task: str,
normalized_state: ndarray,
state_num_bins: int = DEFAULT_STATE_NUM_BINS,
) -> str
Build the π0.5 prompt: instruction + serialized discretized state.
Matches LeRobot's Pi05PrepareStateTokenizerProcessorStep, including the task cleanup (strip, _ → space, newline → space), the exact template and the state values it serializes. The template already ends in a newline, so — unlike π0 — there is no separate newline-appending step.
normalized_state comes from Pi05ForActionPrediction._normalize_state.
create_sinusoidal_pos_embedding ¶
create_sinusoidal_pos_embedding(
time: Tensor,
dimension: int,
min_period: float = 0.004,
max_period: float = 4.0,
device: device = None,
) -> Tensor
Compute a sine/cosine positional embedding for scalar timesteps.
Ref: openpi/models_pytorch/pi0_pytorch.py create_sinusoidal_pos_embedding
discretize_state ¶
discretize_state(
state: ndarray,
*,
num_bins: int = DEFAULT_STATE_NUM_BINS,
) -> ndarray
Discretize a [-1, 1] state into num_bins integer bins.
Byte-for-byte LeRobot's processor_pi05.py::
np.digitize(state_np, bins=np.linspace(-1, 1, 256 + 1)[:-1]) - 1
A state below -1 lands in bin -1, and that negative bin is part of the contract: the checkpoint was trained with " -1" in the state prompt for those dimensions. Clipping it to 0 changes the tokens the model sees, so this must not clip.
bins stays float64, matching LeRobot's default linspace dtype, so boundary values fall on the same side.
Assumes the state is already normalized — see the module docstring.
get_pi05_post_process_func ¶
get_pi05_post_process_func(od_config: OmniDiffusionConfig)
π0.5 returns actions directly; post-processing is identity.
make_att_2d_masks ¶
Build a 2D attention mask from a padding mask and an autoregressive mask.
Ref: openpi/models_pytorch/pi0_pytorch.py make_att_2d_masks
pil_image_to_tensor ¶
PIL → (1, C, H, W) float32 in [-1, 1] (SigLIP normalization).
prepare_attention_masks_4d ¶
Convert (B, S, S) bool masks to (B, 1, S, S) float masks.
True → 0.0 (attend), False → OPENPI_ATTENTION_MASK_VALUE.
resize_with_pad ¶
resize_with_pad(
images: Tensor,
target_height: int,
target_width: int,
mode: str = "bilinear",
) -> Tensor
Resize (B, C, H, W) images to the target shape, preserving aspect ratio with -1 padding on the short side.
Matches openpi image_tools.resize_with_pad_torch — the clamp to [-1, 1] is what lets the padded region blend with SigLIP-normalized pixels without adding signal at the boundary.
resolve_excluded_action_indices ¶
resolve_excluded_action_indices(
exclude_joints: list[str] | None,
action_names: list[str] | None,
) -> list[int]
Map relative_exclude_joints names onto action-vector indices.
Matching is exact name first, then substring (a checkpoint may name the gripper dimension gripper_position while the config just says gripper). Single source of truth for both the config-time validation and the runtime mask in processor_pi05.Pi05RelativeActions.
Returns an empty list when there is nothing to exclude. Raises when a name cannot be resolved — an unresolvable exclusion would otherwise silently become "make this dimension relative too".
tokenize_prompt ¶
tokenize_prompt(
tokenizer,
text: str,
max_token_len: int = DEFAULT_MAX_TOKEN_LEN,
)
Return (input_ids, attention_mask) lists, length exactly max_token_len.
padding="max_length" is what makes the prefix a constant shape: the text segment is always max_token_len tokens regardless of the instruction, so only attention_mask.sum() varies per request.