[https://nvbugs/6481034][test] Trace Kimi gen-only KV transfer timeout - #16918
[https://nvbugs/6481034][test] Trace Kimi gen-only KV transfer timeout#16918chienchunhung wants to merge 7 commits into
Conversation
|
/bot run --disable-fail-fast --stage-list "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4" |
|
PR_Github #62028 [ run ] triggered by Bot. Commit: |
WalkthroughUpdates one perf-sanity waiver and increases the generation worker’s ChangesPerf sanity updates
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
|
/bot run --post-merge --disable-fail-fast --only-multi-gpu-test --stage-list "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4" |
|
PR_Github #62037 [ run ] triggered by Bot. Commit: |
|
PR_Github #62037 Bot args parsing error: CI requested by |
|
/bot run --disable-fail-fast --add-multi-gpu-test --stage-list "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4" |
|
PR_Github #62039 [ run ] triggered by Bot. Commit: |
|
PR_Github #62028 [ run ] completed with state |
|
PR_Github #62039 [ run ] completed with state
|
|
Targeted CI result: the Kimi 0.85 change was not exercised because the perf launcher and pytest selected different shard entries.
This is a deterministic positional-vs-duration shard mismatch, not a failure of the proposed 0.85 setting. A follow-up run should target shard 3 to reach the Kimi pytest case under the current duration mapping; its wrapper will still derive stage-level environment from the positional DeepSeek entry, although the 20-GPU/5-node resource shape is identical. |
|
/bot run --disable-fail-fast --add-multi-gpu-test --stage-list "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3" |
|
/bot run --disable-fail-fast --add-multi-gpu-test --stage-list "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3" |
1 similar comment
|
/bot run --disable-fail-fast --add-multi-gpu-test --stage-list "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3" |
|
PR_Github #62561 [ run ] triggered by Bot. Commit: |
|
PR_Github #62561 [ run ] completed with state
|
|
/bot run --disable-fail-fast --add-multi-gpu-test --stage-list "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3" |
0d14d18 to
2191139
Compare
|
/bot run --disable-fail-fast --add-multi-gpu-test --stage-list "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3" |
|
PR_Github #62618 [ run ] triggered by Bot. Commit: |
|
PR_Github #62618 [ run ] completed with state
|
|
Latest targeted result (#50758) changes the diagnosis:
Therefore the remaining product question is no longer “can this workload complete?” in the latest sample. It is whether #16832 is the actual fix and whether this PR's 0.85 fraction is still necessary. The next targeted run should finally provide the complete per-request trace needed to answer that; an eventual 0.8/fabric-memory A/B can isolate the memory-fraction contribution. |
|
/bot run --disable-fail-fast --add-multi-gpu-test --stage-list "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3" |
|
PR_Github #62854 [ run ] triggered by Bot. Commit: |
|
/bot cancel |
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
GitHub Bot Help
Provide a user friendly way for developers to interact with a Jenkins server. Run See details below for each supported subcommand. Details
Launch build/test pipelines. All previously running jobs will be killed.
kill
Kill all running builds associated with pull request. skip
Skip testing for latest commit on pull request. reuse-pipeline
Reuse a previous pipeline to validate current commit. This action will also kill all currently running builds associated with the pull request. IMPORTANT NOTE: This is dangerous since lack of user care and validation can cause top of tree to break. |
49c05ed to
cf5afef
Compare
|
/bot kill |
|
/bot run --disable-fail-fast --add-multi-gpu-test --stage-list "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3" |
|
PR_Github #62859 [ run ] triggered by Bot. Commit: |
|
PR_Github #62854 [ run ] completed with state |
|
PR_Github #62859 [ run ] completed with state
|
|
Closing because no product or test-list change remains to merge. The exact GB300 Kimi gen-only test was already unwaived by #16717. The targeted validation on a base containing #16832, with GEN free_gpu_memory_fraction=0.8 and kv_transfer_timeout_ms=60000, completed 4,096/4,096 requests successfully; all captured transfers completed below the timeout and the original KV-capacity/transfer-timeout signature did not reproduce. The remaining red stage outcome was a separate post-benchmark GEN srun/completion-sentinel teardown issue and should be tracked independently. Diagnostic tracing and the generic pytest-shard alignment can be split into separate PRs if independently desired. |
Summary
KV_TRANSFER_TRACElifecycle eventsfree_gpu_memory_fraction: 0.8(CTX remains0.6); this PR does not raise the KV-cache fractionleast_durationshard selection so both launch and collect artifacts for the same testValidation result
The targeted run used the Python transceiver with transfer overlap disabled,
kv_transfer_timeout_ms=60000, and GENfree_gpu_memory_fraction: 0.8.The workload itself passed:
num_fitting_reqs=0, insufficient-KV-cache, KV-transfer-timeout, OOM, failedAgentResult, or failed NIXL outcome was observedThe captured transfer latencies were safely below the existing 60-second timeout:
The GEN logs also confirm that #16832's Python-transceiver fabric-memory default was active (
TRTLLM_KVCACHE_POOL_USE_FABRIC_MEMORY=1). The original NVBUG#6481034 KV-capacity/transfer-timeout signature therefore appears fixed by #16832 on the current0.8configuration. This run provides no evidence for raising the memory fraction or transfer timeout.Targeted stage:
GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3Further issue exposed by the run
Although the workload completed and wrote
benchmark_status=Done, the overall stage remained red. All GEN test logic and MPI workers exited cleanly, but the multi-node GENsrundid not finish reaping and therefore never createdgen_server_0.done. The benchmark role waited for that sentinel until one Slurm step exited nonzero;--kill-on-bad-exit=1then cancelled the remaining job.The evidence points to a post-benchmark launcher/teardown problem in the GEN-log completion-sentinel flow introduced by #16717, not a KV-transfer failure. The logs did not capture the exact lingering launcher task, so its identity still needs a teardown-focused diagnostic before assigning a definitive code-level cause.
There is also a separate
ruff-formatfailure injenkins/scripts/perf/submit.py, and this draft currently conflicts with newermain.Current PR scope
Beyond tracing/instrumentation, the diff contains the pytest-shard alignment in
jenkins/scripts/perf/submit.pyand its unit coverage. This was needed because an earlier run had the launcher and pytest selecting different tests/output directories.The net diff contains:
least_durationshard alignment and testsIt contains no net change to KV-cache memory fractions, transfer timeout, or test waivers, and no runtime product fix beyond diagnostics.
Local validation
0.8, CTX0.6python3 -m py_compilepassed for all modified Python filesgit diff --checkpassed