Describe the bug
We attempted to reproduce the Qwen3.8-Flash-Next support described on the AutoModel model-coverage page.
We successfully completed the language-only 100-step H100 workflow using the HellaSwag dataset configuration provided in the published recipe. Training loss decreased from 3.2820 to 1.8538, with a final validation loss of 2.0966.
We observed differences in the source/image composition, loss values, and numerical metrics. These may be expected under different runtime configurations, but a fully specified and pinned reference setup would allow a more direct comparison with the published results.
Steps/Code to reproduce bug
1. AutoModel environment
We used:
The NGC 26.08 container contains embedded AutoModel source commit 922e83e4207c6aaa0ae1654c4621975dffac6849. That revision does not contain the Qwen3.8-Flash-Next implementation, so we mounted commit 9a463234... as a source overlay.
The image provides transformers==5.12.1 and quack-kernels==0.6.1, while the selected source declares transformers==5.15.1 and quack- kernels==0.6.4. The source-overlay setup completed the tested language-only H100 path without issue.
2. H100 training reproduction
We ran the published language-only, full-parameter, 100-step HellaSwag recipe with global batch size 64 on 64 H100 GPUs.
Our run matched the published 64×H100, EP64, gradient-accumulation-1 setup, with using AutoModel's generic torch dispatcher instead of HybridEP. our run completed all 100 optimizer steps and final validation.
-
Published H100: step-0 loss 3.0734, step-99 loss 1.8195, validation loss 2.0610.
-
Local H100: step-0 loss 3.2820, step-99 loss 1.8538, validation loss 2.0966.
3. Checkpoint and numerical validation
We also attempted to reproduce the AutoModel-to-SGLang numerical validation described on the model-coverage page.
The H100 run saved the step-99 model as a native distributed checkpoint. We restored and converted it into an HF-style checkpoint using tools/offline_hf_consolidation.py.
The consolidated checkpoint contained 1,294 language-model tensors, 131 safetensor shards, no visual tensors, no MTP tensors, and language_model_only=true. We checked its tensor names, shapes, dtypes, and selected exact tensor values against the native distributed checkpoint.
SGLang setup:
- Image:
lmsysorg/sglang@sha256:55596775c4c77306d19f42f9991add59c53b7ec1cce83b11631486695779b3c2
- Source commit:
593134d17a6eb0d0fc5f71a970cd2e9dc8e26e8b
We performed two comparisons:
- Released checkpoint: AutoModel and SGLang loaded the same released checkpoint.
- Trained checkpoint: AutoModel restored the native step-99 checkpoint, while SGLang loaded a metadata-compatible view of its consolidated HF representation with the same weight payload. Within each comparison, both runtimes received the same fixed raw token IDs. The local topologies were AutoModel FSDP2/EP64/TP1 on 64 H100 GPUs and SGLang TP8/EP8 on 8 H100 GPUs.
Released checkpoint:
- Ordinary 64: cosine
0.99601894, relative L2 0.09052518, top-1 same.
- EOS-reset 64: cosine
0.98668198, relative L2 0.16775986, top-1 same.
- QSA 4,096: cosine
0.99690736, relative L2 0.07905925, top-1 same.
Trained step-99 checkpoint:
- Ordinary 64: cosine
0.99810641, relative L2 0.06166070, top-1 same.
- EOS-reset 64: cosine
0.99333476, relative L2 0.11550583, top-1 same.
- QSA 4,096: cosine
0.99632406, relative L2 0.08700404, top-1 different with near-tied logits.
The published numerical validation applies to the released checkpoint and reports final-logit cosine 0.99952, relative L2 0.03175, and identical top-1/top-5 tokens. It also reports first-four-layer hidden-state cosine 0.9998–1.0000, Engram relative L2 below 0.0063, and QSA projected-output relative L2 0.0041.
Expected behavior
It would be helpful to clarify:
- What source commit, image were used for the published 100-step H100 result?
- What SGLang image/revision/config/fixed inputs used for the published numerical validation?
- When model support is available in source before a matching NGC image is released, is building a custom image from that source commit's Dockerfile the recommended way to reproduce the example?
Describe the bug
We attempted to reproduce the Qwen3.8-Flash-Next support described on the AutoModel model-coverage page.
We successfully completed the language-only 100-step H100 workflow using the HellaSwag dataset configuration provided in the published recipe. Training loss decreased from
3.2820to1.8538, with a final validation loss of2.0966.We observed differences in the source/image composition, loss values, and numerical metrics. These may be expected under different runtime configurations, but a fully specified and pinned reference setup would allow a more direct comparison with the published results.
Steps/Code to reproduce bug
1. AutoModel environment
We used:
9a4632347323fbfee9ba1ae1a323669ef74fc5fbnvcr.io/nvidia/nemo-automodel@sha256:c21c0c3b82f68fcccfded4bf9fe11b2d53df1c9a969e9e883a1562ffca6294f6The NGC 26.08 container contains embedded AutoModel source commit
922e83e4207c6aaa0ae1654c4621975dffac6849. That revision does not contain the Qwen3.8-Flash-Next implementation, so we mounted commit9a463234...as a source overlay.The image provides
transformers==5.12.1andquack-kernels==0.6.1, while the selected source declarestransformers==5.15.1andquack- kernels==0.6.4. The source-overlay setup completed the tested language-only H100 path without issue.2. H100 training reproduction
We ran the published language-only, full-parameter, 100-step HellaSwag recipe with global batch size 64 on 64 H100 GPUs.
Our run matched the published 64×H100, EP64, gradient-accumulation-1 setup, with using AutoModel's generic
torchdispatcher instead of HybridEP. our run completed all 100 optimizer steps and final validation.Published H100: step-0 loss
3.0734, step-99 loss1.8195, validation loss2.0610.Local H100: step-0 loss
3.2820, step-99 loss1.8538, validation loss2.0966.3. Checkpoint and numerical validation
We also attempted to reproduce the AutoModel-to-SGLang numerical validation described on the model-coverage page.
The H100 run saved the step-99 model as a native distributed checkpoint. We restored and converted it into an HF-style checkpoint using
tools/offline_hf_consolidation.py.The consolidated checkpoint contained 1,294 language-model tensors, 131 safetensor shards, no visual tensors, no MTP tensors, and
language_model_only=true. We checked its tensor names, shapes, dtypes, and selected exact tensor values against the native distributed checkpoint.SGLang setup:
lmsysorg/sglang@sha256:55596775c4c77306d19f42f9991add59c53b7ec1cce83b11631486695779b3c2593134d17a6eb0d0fc5f71a970cd2e9dc8e26e8bWe performed two comparisons:
Released checkpoint:
0.99601894, relative L20.09052518, top-1 same.0.98668198, relative L20.16775986, top-1 same.0.99690736, relative L20.07905925, top-1 same.Trained step-99 checkpoint:
0.99810641, relative L20.06166070, top-1 same.0.99333476, relative L20.11550583, top-1 same.0.99632406, relative L20.08700404, top-1 different with near-tied logits.The published numerical validation applies to the released checkpoint and reports final-logit cosine
0.99952, relative L20.03175, and identical top-1/top-5 tokens. It also reports first-four-layer hidden-state cosine0.9998–1.0000, Engram relative L2 below0.0063, and QSA projected-output relative L20.0041.Expected behavior
It would be helpful to clarify: