Benchmarking¶
Benchmarking defines reproducible workloads, metrics, warm-up, measurement, and result metadata for comparing vLLM-Omni performance.
Candidate invariants¶
BENCH-INV-001: Results are reproducible¶
Rule: Results MUST identify the model, revision, commit, hardware, configuration, workload, concurrency, warm-up, and measurement window.
BENCH-INV-002: Correctness precedes performance¶
Rule: A benchmark MUST validate successful and semantically acceptable outputs before using timings as performance evidence.
BENCH-INV-003: Metric definitions are explicit¶
Rule: Metrics MUST state their unit, population, and aggregation.
Safe-change guide¶
Test workload parsing and metric calculations with deterministic fixtures. Do not compare materially different workloads or runtime settings.