Skip to content

Questions about reproducing the Qwen3.8-Flash-Next example #4009

Description

@rouchenzi

Describe the bug

We attempted to reproduce the Qwen3.8-Flash-Next support described on the AutoModel model-coverage page.

We successfully completed the language-only 100-step H100 workflow using the HellaSwag dataset configuration provided in the published recipe. Training loss decreased from 3.2820 to 1.8538, with a final validation loss of 2.0966.

We observed differences in the source/image composition, loss values, and numerical metrics. These may be expected under different runtime configurations, but a fully specified and pinned reference setup would allow a more direct comparison with the published results.

Steps/Code to reproduce bug

1. AutoModel environment

We used:

The NGC 26.08 container contains embedded AutoModel source commit 922e83e4207c6aaa0ae1654c4621975dffac6849. That revision does not contain the Qwen3.8-Flash-Next implementation, so we mounted commit 9a463234... as a source overlay.

The image provides transformers==5.12.1 and quack-kernels==0.6.1, while the selected source declares transformers==5.15.1 and quack- kernels==0.6.4. The source-overlay setup completed the tested language-only H100 path without issue.

2. H100 training reproduction

We ran the published language-only, full-parameter, 100-step HellaSwag recipe with global batch size 64 on 64 H100 GPUs.

Our run matched the published 64×H100, EP64, gradient-accumulation-1 setup, with using AutoModel's generic torch dispatcher instead of HybridEP. our run completed all 100 optimizer steps and final validation.

  • Published H100: step-0 loss 3.0734, step-99 loss 1.8195, validation loss 2.0610.

  • Local H100: step-0 loss 3.2820, step-99 loss 1.8538, validation loss 2.0966.

3. Checkpoint and numerical validation

We also attempted to reproduce the AutoModel-to-SGLang numerical validation described on the model-coverage page.

The H100 run saved the step-99 model as a native distributed checkpoint. We restored and converted it into an HF-style checkpoint using tools/offline_hf_consolidation.py.

The consolidated checkpoint contained 1,294 language-model tensors, 131 safetensor shards, no visual tensors, no MTP tensors, and language_model_only=true. We checked its tensor names, shapes, dtypes, and selected exact tensor values against the native distributed checkpoint.

SGLang setup:

  • Image: lmsysorg/sglang@sha256:55596775c4c77306d19f42f9991add59c53b7ec1cce83b11631486695779b3c2
  • Source commit: 593134d17a6eb0d0fc5f71a970cd2e9dc8e26e8b

We performed two comparisons:

  1. Released checkpoint: AutoModel and SGLang loaded the same released checkpoint.
  2. Trained checkpoint: AutoModel restored the native step-99 checkpoint, while SGLang loaded a metadata-compatible view of its consolidated HF representation with the same weight payload. Within each comparison, both runtimes received the same fixed raw token IDs. The local topologies were AutoModel FSDP2/EP64/TP1 on 64 H100 GPUs and SGLang TP8/EP8 on 8 H100 GPUs.

Released checkpoint:

  • Ordinary 64: cosine 0.99601894, relative L2 0.09052518, top-1 same.
  • EOS-reset 64: cosine 0.98668198, relative L2 0.16775986, top-1 same.
  • QSA 4,096: cosine 0.99690736, relative L2 0.07905925, top-1 same.

Trained step-99 checkpoint:

  • Ordinary 64: cosine 0.99810641, relative L2 0.06166070, top-1 same.
  • EOS-reset 64: cosine 0.99333476, relative L2 0.11550583, top-1 same.
  • QSA 4,096: cosine 0.99632406, relative L2 0.08700404, top-1 different with near-tied logits.

The published numerical validation applies to the released checkpoint and reports final-logit cosine 0.99952, relative L2 0.03175, and identical top-1/top-5 tokens. It also reports first-four-layer hidden-state cosine 0.9998–1.0000, Engram relative L2 below 0.0063, and QSA projected-output relative L2 0.0041.

Expected behavior

It would be helpful to clarify:

  1. What source commit, image were used for the published 100-step H100 result?
  2. What SGLang image/revision/config/fixed inputs used for the published numerical validation?
  3. When model support is available in source before a matching NGC image is released, is building a custom image from that source commit's Dockerfile the recommended way to reproduce the example?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions