Skip to content

[MaxText] Fix the MLPerf parallelism, precision and config_filename mllog disclosure for Lineage runs - #5519

Merged
copybara-service[bot] merged 1 commit into
mainfrom
test_992381189
Oct 3, 2026
Merged

copybara-service[bot] merged 1 commit into
mainfrom
test_992381189

Conversation

@copybara-service

Copy link
Copy Markdown
Contributor

[MaxText] Fix the MLPerf parallelism, precision and config_filename mllog disclosure for Lineage runs

mllog_utils.init_print derived the v6.1 disclosure keys from the named
ici_*_parallelism and quantization fields. The Lineage recipe
(deepseek3-671b-lineage.yml) uses neither: it runs on a physical
[dcn, x, y, z, core] mesh with explicit logical_axis_rules and quantizes
through lineage_quantization. The 2026-10-01 Lineage runs therefore logged
expert_parallelism=1, bfloat16 for linear/comm, and every run logged the
placeholder config_filename=config.yml.

For use_lineage runs:

  • expert_parallelism is the mesh size of the exp rule
    (x * y * core = 32).
  • tensor_parallelism is the mesh size of the activation_length rule
    (the TensorCore pair, 2). Lineage's MLA up/out projections are head-sharded
    across it, and the splash kernel sees the full sequence (Megatron TP + SP),
    so this is not context parallelism.
  • precision falls back to lineage_quantization. fp8_full quantizes
    the routed-expert GMMs and the EP token all-gather
    (QuantConfig.routed_experts, dsv3_sparse_layer.ubatch_dispatch), so both
    linear and comm log fp8.

For all runs, config_filename is <model_name>.yml.

Non-Lineage parallelism and precision are unchanged.

@codecov

codecov Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.90909% with 2 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/maxtext/utils/mllog_utils.py 90.90% 1 Missing and 1 partial ⚠️

📢 Thoughts on this report? Let us know!

…llog disclosure for Lineage runs

`mllog_utils.init_print` derived the v6.1 disclosure keys from the named
`ici_*_parallelism` and `quantization` fields. The Lineage recipe
(`deepseek3-671b-lineage.yml`) uses neither: it runs on a physical
`[dcn, x, y, z, core]` mesh with explicit `logical_axis_rules` and quantizes
through `lineage_quantization`. The 2026-10-01 Lineage runs therefore logged
`expert_parallelism=1`, `bfloat16` for linear/comm, and every run logged the
placeholder `config_filename=config.yml`.

For `use_lineage` runs:

- **expert_parallelism** is the mesh size of the `exp` rule
  (`x * y * core` = 32).
- **tensor_parallelism** is the mesh size of the `activation_length` rule
  (the TensorCore pair, 2). Lineage's MLA up/out projections are head-sharded
  across it, and the splash kernel sees the full sequence (Megatron TP + SP),
  so this is not context parallelism.
- **precision** falls back to `lineage_quantization`. `fp8_full` quantizes
  the routed-expert GMMs and the EP token all-gather
  (`QuantConfig.routed_experts`, `dsv3_sparse_layer.ubatch_dispatch`), so both
  linear and comm log `fp8`.

For all runs, **config_filename** is `<model_name>.yml`.

Non-Lineage parallelism and precision are unchanged.

PiperOrigin-RevId: 992750871
@copybara-service
copybara-service Bot merged commit e3f76fb into main Oct 3, 2026
4 checks passed
@copybara-service
copybara-service Bot deleted the test_992381189 branch October 3, 2026 07:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant