Multi-Level Automated Testing System Overview¶
Document Overview¶
This testing system aims to build a complete, efficient, and well-structured quality assurance framework for the development, integration, and release of model services. It draws on the concept of the test pyramid from modern software engineering, progressively expanding testing activities from basic code logic verification to complex end-to-end (E2E) functionality, performance, accuracy, and even long-term stability validation.
Through five levels (L1-L5) and common (Common) specifications, the system clarifies the testing objectives, scope, execution frequency, and required resources for different development stages (e.g., each commit, PR merge, daily build, pre-release). This ensures that models meet high standards for functionality, performance, and reliability across various deployment scenarios (online serving and offline inference).
| Level | Scope & Focus | Model Coverage Strategy | Module Strategy | Tags | Time Cost | Test Dir | Doc | Frequency | Hardware |
|---|---|---|---|---|---|---|---|---|---|
| Common | Contribution Guideline & PR checklist | / | / | / | / | / | PR Checklist | / | / |
| CI Failure Description | / | / | / | / | / | CI Failures | / | / | |
| L1 (Unit & Logic) | Unit tests for components like entrypoints, models | / | / | core_model and cpu | <15min | tests/{component}/… mirroring vllm_omni/{component}/(e.g. tests/diffusion/, tests/engine/, tests/entrypoints/)Do not add new top-level tests/ dirs unrelated to a component. | L1 & L2 Section 1 L1&L2: Purpose, Test Content, Directory Location, Example | PR with ready label (also can run locally) | CPU |
| L2 (E2E across models & GPU-required UT) | Online (basic deployment scenarios): dummy, normal inference function (output format, stream), some instance startup UT | Key models + online basic scenarios; request success, non-empty output, format match (no Whisper/accuracy) | Key features and modules that require launching instances (prefer random weights) | core_model and hardware_test(H100, L4, etc.) and omni/tts/diffusion | Model E2E:tests/e2e/online_serving/test_{model}.pytests/e2e/offline_inference/test_{model}.pyFeature integration: tests/e2e/features/<feature>/(e.g. fullduplex/, custom_pipeline/, rlhf_test/)Component / interface Test (GPU): tests/{component}/…, tests/entrypoints/… | L1 & L2 L1&L2: Purpose, Test Content, Directory Location, Example | PR with ready label | GPU | |
| L3 (Important Perf & Integration & Accuracy) | Online & Offline (multiple deployment scenarios): real model, normal inference function, normal accuracy | Key models + key online/offline scenarios; real weights, Whisper/similarity, preset voice gender, basic accuracy | Key features and modules that require launching instances (using real weights) | advanced_model and hardware_test(H100, L4, etc.) and omni/tts/diffusion | <30min | Model E2E:tests/e2e/online_serving/test_{model}.pytests/e2e/offline_inference/test_{model}.pytests/e2e/accuracy/Feature integration: tests/e2e/features/<feature>/Component / interface Test: tests/{component}/…, tests/entrypoints/… | L3 L3: Purpose, Test Content, Directory Location, Example | PR Merged (Also run L1&L2 Tests) | GPU |
| L4 (Perf & Integration & Accuracy) | Online: full functional scenarios + performance test + doc test + accuracy test | Key models: function, performance, accuracy, and doc testing | Other features and modules that require launching instances (using real weights) | full_model and hardware_test(H100, L4, etc.) and omni/tts/diffusion | <3 hour | Model E2E:tests/e2e/online_serving/test_{model}_expansion.pytests/e2e/offline_inference/test_{model}_expansion.pytests/e2e/accuracy/test_{model}.pyFeature integration: tests/e2e/features/<feature>/Component / interface Test: tests/{component}/…, tests/entrypoints/…Performance: tests/dfx/perf/tests/Doc examples: tests/examples/online_serving/, tests/examples/offline_inference/ | L4 L4: Purpose, Test Content, Directory Location, Example | Nightly / Days before Release (additional GPU SKUs) | GPU |
| L5 (Stability & Reliability & selected Perf) | Online: long-term stability + reliability; coverage test | Long-term stability and reliability testing for key models Non-critical scenario performance and accuracy testing for key models Other models: function and doc testing | / | slow and hardware_test(H100, L4, etc.) and omni/tts/diffusion | Depends on reality | Stability:tests/dfx/stability/tests/Reliability: tests/dfx/reliability/test_reliability_{model}.py | L5 L5: Purpose, Test Content, Directory Location, Example | Weekly / Days before Release | GPU |
Test Dir placement (summary): component / unit under tests/{component}/ mirroring vllm_omni/; model E2E under tests/e2e/online_serving/, tests/e2e/offline_inference/, and tests/e2e/accuracy/; feature integration under tests/e2e/features/<feature>/; doc example tests under tests/examples/online_serving/ and tests/examples/offline_inference/; performance under tests/dfx/perf/; stability under tests/dfx/stability/; reliability under tests/dfx/reliability/. Do not add new top-level directories under tests/ that are unrelated to a vllm_omni component (or to the established e2e / dfx / helpers / examples / buildkite layout).
For per-level test authoring (markers, examples), see Test Writing Guide.
Common Specifications¶
Before entering specific testing levels, the project establishes two common specifications aimed at standardizing the development process and quickly locating issues.
- PR Checklist (
.github/PULL_REQUEST_TEMPLATE.md): This template defines the self-check items that must be completed before submitting a code review (Pull Request). It ensures that each code change meets basic requirements such as code style, dependency updates, and documentation synchronization before entering the automated testing pipeline, serving as the first manual line of defense for quality assurance. - CI Failure Explanation (CI Failures): This document archives and explains common failure patterns in the Continuous Integration (CI) pipeline, error log interpretation, and preliminary troubleshooting steps. It helps developers and testers quickly diagnose the causes of automated test failures, improving problem-solving efficiency.
Notes¶
L4 cadence (nightly and pre-release)¶
CUDA L4 (full_model) jobs are defined once in .buildkite/cuda/test-nightly.yml. That file is reused at two cadences:
- Nightly: scheduled
mainbuilds withNIGHTLY=1(and PR labels such asnightly-test/omni-test). Default fleet is typically H100. - Pre-release: the same L4 suite runs again on extra GPU SKUs before a release (for example B200 via
MIRROR_HW=b200), sofull_modelcoverage is not limited to the nightly machine type.
AMD has an experimental, non-blocking full_model nightly vertical slice in .buildkite/amd/test-amd-nightly.yml. It selects ROCm Qwen3-Omni function-expansion, accuracy, documentation-example, and AITER-on smoke coverage on the MI300 pool. AMD-only PR validation uses amd-test to admit the external AMD pipeline and nightly-test to select this suite; scheduled main builds select it with NIGHTLY=1. Keep it non-blocking while the lane accumulates stability and completion-time evidence.
How count-form mirror_hardwares and MIRROR_HW pick a SKU: CI Settings.
L2 / L3 diff-aware CI (CUDA)¶
On CUDA L2 and L3, E2E Test Buildkite jobs may be omitted at pipeline upload when the PR diff does not touch their path prefixes; other groups still always upload. Full two-layer mechanics and YAML examples: CI Settings — Diff-aware CI (step filtering).
Test helper environment variables¶
Some shared helpers under tests/helpers/ honor optional environment variables for local debugging. These are not set in CI by default.
| Variable | Accepted values | Description |
|---|---|---|
VLLM_OMNI_KEEP_REQUEST_MEDIA | 1, true, yes (case-insensitive) | When enabled, temporary WAV files created by tests.helpers.media.convert_audio_bytes_to_text are not deleted when the pytest process exits. By default, each call writes a unique file under the system temp directory via tempfile.mkstemp and registers atexit cleanup. Use this when debugging audio output validation (Whisper transcription, keyword checks, text–audio similarity). The saved path is logged as audio data is saved: <path>. |
Example (Linux / macOS):
export VLLM_OMNI_KEEP_REQUEST_MEDIA=1
pytest -s -v tests/e2e/online_serving/test_qwen3_omni.py -k test_mix_to_text_audio
Example (Windows PowerShell):
$env:VLLM_OMNI_KEEP_REQUEST_MEDIA = "1"
pytest -s -v tests/e2e/online_serving/test_qwen3_omni.py -k test_mix_to_text_audio
Summary¶
This multi-level testing system achieves continuous, progressive validation of model service quality by tightly integrating testing activities with the development workflow (commit, review, merge, release). From rapid unit testing to comprehensive end-to-end testing, and further to in-depth performance, stability, and reliability verification, each level has clear objectives, collectively building a robust quality protection net. By following this system, teams can deliver high-quality, highly reliable model services more efficiently.