Stage runtime and replica lifecycle¶
The stage runtime turns a logical stage definition into ready local or distributed replicas and provides the lifecycle boundary used by the orchestrator.
Contract status¶
This document describes the current StageRuntime, StagePool, stage-client, and stage-process layout. The unified LLM/diffusion direction in #5441 remains in flight and is not presented as current behavior.
Ownership boundary¶
This document owns local and distributed stage placement, startup and readiness, one pool per logical stage, replica identity and membership, selection and affinity, stage-client/process lifecycle, liveness, draining, and shutdown.
Local EngineCore initialization and multi-API serving use StageRuntime.launch_stage_engines(num_api_servers=1) as their common launch entry. It resolves plans when not supplied by the runtime's device-group scheduler, coordinates device locks and environment overlays, starts engines, and rolls back resources on startup or attachment failure. Single-API clients attach locally after readiness and take ownership of the engine resources; there is no additional API subprocess. Diffusion and remote attachment retain their backend-specific initialization.
For local multi-API serving, the parent runtime owns and launches each stage engine once. It supplies a distinct input/output channel for every API frontend and stage replica; frontend workers attach clients to those engines without acquiring stage-process ownership. Readiness observation and shutdown remain parent-owned. This topology currently excludes diffusion and remote/headless stages, intra-stage data parallelism, and Ray backends.
It does not own cross-stage request policy, model scheduler policy, connector transfer semantics, configuration precedence, or public error rendering.
Candidate invariants¶
These identifiers are proposals while the document is draft.
STAGE-INV-001: The runtime owns replica lifecycle¶
Rule: Orchestration MUST acquire stage capacity through the stage runtime and MUST NOT construct or retire backend replicas directly.
STAGE-INV-100: Affinity uses stable replica identity¶
Rule: Follow-up operations for a request with replica affinity MUST resolve the same logical replica until the request terminates or the replica is declared lost.
STAGE-INV-200: Shutdown is idempotent¶
Rule: Repeated shutdown or cleanup signals MUST NOT leak stage clients, process managers, or membership registrations.
Invariant namespace¶
STAGE-INV reserves 001-099 for placement and ownership, 100-199 for replica identity/readiness/affinity, 200-299 for liveness loss, draining and cleanup, and 300-399 for local/distributed and upstream compatibility. Numbers become append-only after normative promotion.
Safe-change guide¶
Exercise local and distributed startup, readiness, replica selection, affinity, membership changes, abort routing, process failure, draining, and repeated shutdown. Cross-stage policy changes belong in engine_orchestration.md.
For multi-API changes, additionally exercise per-frontend channel assignment, shared-engine startup handshakes, frontend failure observation, and parent cleanup without duplicate stage shutdown.
Promotion gate¶
- Reconcile terminology and ownership after #5441 reaches a final state.
- Add or identify a dedicated StageRuntime/StagePool lifecycle suite; current evidence is distributed across the cited tests.
- Demonstrate affinity preservation for update, interaction, and abort paths.
- Demonstrate membership loss and repeated shutdown without leaked clients or processes.
- Obtain approval from a technical owner and an independent validation reviewer.