[MaxText] Fix the MLPerf parallelism, precision and config_filename mllog disclosure for Lineage runs - #5519
Merged
Merged
Conversation
copybara-service
Bot
requested review from
A9isha,
NuojCheng,
RissyRan,
SurbhiJainUSC,
abhinavclemson,
aireenmei,
bvandermoon,
darisoy,
dipannita08,
gagika,
gobbleturk,
hengtaoguo,
huytransformer,
igorts-git,
jiangjy1982,
khatwanimohit,
richjames0,
shralex,
shuningjin,
vipannalla and
xibinliu
as code owners
October 2, 2026 22:41
copybara-service
Bot
force-pushed
the
test_992381189
branch
from
October 2, 2026 22:50
37d5eef to
e1b0605
Compare
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
copybara-service
Bot
force-pushed
the
test_992381189
branch
from
October 3, 2026 06:58
e1b0605 to
1ad8f8c
Compare
…llog disclosure for Lineage runs `mllog_utils.init_print` derived the v6.1 disclosure keys from the named `ici_*_parallelism` and `quantization` fields. The Lineage recipe (`deepseek3-671b-lineage.yml`) uses neither: it runs on a physical `[dcn, x, y, z, core]` mesh with explicit `logical_axis_rules` and quantizes through `lineage_quantization`. The 2026-10-01 Lineage runs therefore logged `expert_parallelism=1`, `bfloat16` for linear/comm, and every run logged the placeholder `config_filename=config.yml`. For `use_lineage` runs: - **expert_parallelism** is the mesh size of the `exp` rule (`x * y * core` = 32). - **tensor_parallelism** is the mesh size of the `activation_length` rule (the TensorCore pair, 2). Lineage's MLA up/out projections are head-sharded across it, and the splash kernel sees the full sequence (Megatron TP + SP), so this is not context parallelism. - **precision** falls back to `lineage_quantization`. `fp8_full` quantizes the routed-expert GMMs and the EP token all-gather (`QuantConfig.routed_experts`, `dsv3_sparse_layer.ubatch_dispatch`), so both linear and comm log `fp8`. For all runs, **config_filename** is `<model_name>.yml`. Non-Lineage parallelism and precision are unchanged. PiperOrigin-RevId: 992750871
copybara-service
Bot
force-pushed
the
test_992381189
branch
from
October 3, 2026 07:43
1ad8f8c to
e3f76fb
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
[MaxText] Fix the MLPerf parallelism, precision and config_filename mllog disclosure for Lineage runs
mllog_utils.init_printderived the v6.1 disclosure keys from the namedici_*_parallelismandquantizationfields. The Lineage recipe(
deepseek3-671b-lineage.yml) uses neither: it runs on a physical[dcn, x, y, z, core]mesh with explicitlogical_axis_rulesand quantizesthrough
lineage_quantization. The 2026-10-01 Lineage runs therefore loggedexpert_parallelism=1,bfloat16for linear/comm, and every run logged theplaceholder
config_filename=config.yml.For
use_lineageruns:exprule(
x * y * core= 32).activation_lengthrule(the TensorCore pair, 2). Lineage's MLA up/out projections are head-sharded
across it, and the splash kernel sees the full sequence (Megatron TP + SP),
so this is not context parallelism.
lineage_quantization.fp8_fullquantizesthe routed-expert GMMs and the EP token all-gather
(
QuantConfig.routed_experts,dsv3_sparse_layer.ubatch_dispatch), so bothlinear and comm log
fp8.For all runs, config_filename is
<model_name>.yml.Non-Lineage parallelism and precision are unchanged.