O2iginal's picture
Upload LOG_NODE_RANK_0.log to based-distill56l-dclm10b-s512-step394-mamba_hy-2.9b-A0-hd64-ng6-msdim320-me1-bs1024-sl32768
35ab7cd verified
Raw
History Blame
224 kB
using world size: 8, data-parallel size: 2, context-parallel size: 2, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 2, encoder-tensor-model-parallel size: 2, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0
WARNING: overriding default arguments for tokenizer_type:GPT2BPETokenizer with tokenizer_type:HuggingFaceTokenizer
Number of virtual stages per pipeline stage: None
accumulate and all-reduce gradients in fp32 for bfloat16 data type.
using torch.bfloat16 for parameters ...
------------------------ arguments ------------------------
account_for_embedding_in_pipeline_split ......... False
account_for_loss_in_pipeline_split .............. False
accumulate_allreduce_grads_in_fp32 .............. True
adam_beta1 ...................................... 0.9
adam_beta2 ...................................... 0.999
adam_eps ........................................ 1e-08
add_bias_linear ................................. False
add_position_embedding .......................... False
add_qkv_bias .................................... True
adlr_autoresume ................................. False
adlr_autoresume_interval ........................ 1000
align_grad_reduce ............................... True
align_param_gather .............................. False
app_tag_run_name ................................ None
app_tag_run_version ............................. 0.0.0
apply_layernorm_1p .............................. False
apply_query_key_layer_scaling ................... False
apply_residual_connection_post_layernorm ........ False
apply_rope_fusion ............................... True
async_save ...................................... None
async_tensor_model_parallel_allreduce ........... True
attention_backend ............................... AttnBackend.auto
attention_dropout ............................... 0.0
attention_softmax_in_fp32 ....................... False
attn_output_gate ................................ None
attn_token_shift ................................ None
auto_detect_ckpt_format ......................... False
barrier_with_L1_time ............................ True
bert_binary_head ................................ True
bert_embedder_type .............................. megatron
bert_load ....................................... None
bf16 ............................................ True
bias_dropout_fusion ............................. True
bias_gelu_fusion ................................ False
bias_swiglu_fusion .............................. True
biencoder_projection_dim ........................ 0
biencoder_shared_query_context_model ............ False
block_data_path ................................. None
calc_ft_timeouts ................................ False
calculate_per_token_loss ........................ False
check_for_large_grads ........................... False
check_for_nan_in_loss_and_grad .................. True
check_for_spiky_loss ............................ False
check_weight_hash_across_dp_replicas_interval ... None
ckpt_assume_constant_structure .................. False
ckpt_convert_format ............................. None
ckpt_convert_save ............................... None
ckpt_convert_update_legacy_dist_opt_format ...... False
ckpt_format ..................................... torch
ckpt_fully_parallel_load ........................ False
ckpt_fully_parallel_save ........................ True
ckpt_fully_parallel_save_deprecated ............. False
ckpt_step ....................................... None
classes_fraction ................................ 1.0
clip_grad ....................................... 0.5
clone_scatter_output_in_embedding ............... True
config_logger_dir ...............................
consumed_train_samples .......................... 0
consumed_valid_samples .......................... 0
context_parallel_size ........................... 2
cp_comm_type .................................... ['p2p']
create_attention_mask_in_dataloader ............. False
cross_entropy_fusion_impl ....................... native
cross_entropy_loss_fusion ....................... False
cuda_graph_scope ................................ full
cuda_graph_warmup_steps ......................... 3
data_args_path .................................. None
data_cache_path ................................. /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache
data_parallel_random_init ....................... False
data_parallel_sharding_strategy ................. no_shard
data_parallel_size .............................. 2
data_path ....................................... ['/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache/datasets/huggingface/Teaven/combine_2B_0908/binidx/yulan_mini']
data_per_class_fraction ......................... 1.0
data_sharding ................................... True
dataloader_type ................................. single
ddp_average_in_collective ....................... False
ddp_bucket_size ................................. None
ddp_num_buckets ................................. None
ddp_pad_buckets_for_high_nccl_busbw ............. False
decoder_first_pipeline_num_layers ............... None
decoder_last_pipeline_num_layers ................ None
decoder_num_layers .............................. None
decoder_seq_length .............................. None
decoupled_lr .................................... None
decoupled_min_lr ................................ None
decrease_batch_size_if_needed ................... False
defer_embedding_wgrad_compute ................... False
deprecated_use_mcore_models ..................... True
deterministic_mode .............................. False
dino_bottleneck_size ............................ 256
dino_freeze_last_layer .......................... 1
dino_head_hidden_size ........................... 2048
dino_local_crops_number ......................... 10
dino_local_img_size ............................. 96
dino_norm_last_layer ............................ False
dino_teacher_temp ............................... 0.07
dino_warmup_teacher_temp ........................ 0.04
dino_warmup_teacher_temp_epochs ................. 30
disable_bf16_reduced_precision_matmul ........... False
disable_mamba_mem_eff_path ...................... False
disable_straggler_on_startup .................... False
dist_ckpt_format_deprecated ..................... None
dist_ckpt_strictness ............................ assume_ok_unexpected
distribute_saved_activations .................... False
distributed_backend ............................. nccl
distributed_timeout_minutes ..................... 10
emb_deviation_loss_coeff ........................ 0
emb_deviation_type .............................. None
embedding_path .................................. None
empty_unused_memory_level ....................... 0
enable_cuda_graph ............................... False
enable_ft_package ............................... False
enable_gloo_process_groups ...................... True
enable_msc ...................................... True
enable_one_logger ............................... True
encoder_num_layers .............................. 112
encoder_pipeline_model_parallel_size ............ 0
encoder_seq_length .............................. 32768
encoder_tensor_model_parallel_size .............. 2
end_weight_decay ................................ 0.1
eod_mask_loss ................................... False
error_injection_rate ............................ 0
error_injection_type ............................ transient_error
eval_interval ................................... 1000
eval_iters ...................................... 10
evidence_data_path .............................. None
exit_duration_in_mins ........................... None
exit_interval ................................... None
exit_on_missing_checkpoint ...................... False
exit_signal_handler ............................. False
exp_avg_dtype ................................... torch.float32
exp_avg_sq_dtype ................................ torch.float32
expert_model_parallel_size ...................... 1
expert_tensor_parallel_size ..................... 2
external_cuda_graph ............................. False
ffn_hidden_size ................................. 4800
ffn_token_shift ................................. None
finetune ........................................ False
first_last_layers_bf16 .......................... False
flash_decode .................................... False
fp16 ............................................ False
fp16_lm_cross_entropy ........................... False
fp32_residual_connection ........................ False
fp8 ............................................. None
fp8_amax_compute_algo ........................... most_recent
fp8_amax_history_len ............................ 1
fp8_interval .................................... 1
fp8_margin ...................................... 0
fp8_param_gather ................................ False
fp8_recipe ...................................... delayed
fp8_wgrad ....................................... True
freeze_non_mamba ................................ False
geglu ........................................... False
global_batch_size ............................... 1024
grad_reduce_in_bf16 ............................. False
gradient_accumulation_fusion .................... True
gradient_reduce_div_fusion ...................... True
group_query_attention ........................... True
head_lr_mult .................................... 1.0
heterogeneous_layers_config_encoded_json ........ None
heterogeneous_layers_config_path ................ None
hidden_dropout .................................. 0.0
hidden_size ..................................... 1920
hierarchical_context_parallel_sizes ............. None
hybrid_attention_ratio .......................... 0.0625
hybrid_mlp_ratio ................................ 0.5
hybrid_override_pattern ......................... *-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-
hysteresis ...................................... 2
ict_head_size ................................... None
ict_load ........................................ None
img_h ........................................... 224
img_w ........................................... 224
increase_log_level_interval ..................... 1000
increase_log_level_iters ........................ 5
indexer_batch_size .............................. 128
indexer_log_interval ............................ 1000
inference_batch_times_seqlen_threshold .......... -1
inference_dynamic_batching ...................... False
inference_dynamic_batching_buffer_guaranteed_fraction 0.2
inference_dynamic_batching_buffer_overflow_factor None
inference_dynamic_batching_buffer_size_gb ....... 40.0
inference_dynamic_batching_chunk_size ........... 256
inference_dynamic_batching_max_requests_override None
inference_dynamic_batching_max_tokens_override .. None
inference_max_batch_size ........................ 8
inference_max_seq_length ........................ 2560
inference_rng_tracker ........................... False
init_method_std ................................. 0.02
init_method_xavier_uniform ...................... False
init_model_with_meta_device ..................... False
initial_loss_scale .............................. 4294967296
is_hybrid_model ................................. False
iter_per_epoch .................................. 1250
iterations_to_skip .............................. []
keep_fp8_transpose_cache_when_using_custom_fsdp . False
kv_channels ..................................... 64
kv_lora_rank .................................... 32
lazy_mpu_init ................................... None
load ............................................ /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/RADLADS-paper/out/L56-D1920-qwen_mamba2_qwen2-e1-i1920-s320-hd64-gn6-A0-S512--step1-dclm10b/rwkv-394-hf-A7-0_8_16_24_32_40_48/megatron-pp1-tp2
local_rank ...................................... 0
log_interval .................................... 1
log_layer_hidden_states ......................... []
log_loss_scale_to_tensorboard ................... True
log_memory_to_tensorboard ....................... True
log_num_zeros_in_grad ........................... False
log_params_norm ................................. True
log_progress .................................... False
log_straggler ................................... False
log_throughput .................................. True
log_timers_to_tensorboard ....................... True
log_validation_ppl_to_tensorboard ............... False
log_world_size_to_tensorboard ................... False
logging_level ................................... None
loss_scale ...................................... None
loss_scale_window ............................... 1000
lr .............................................. 2e-05
lr_decay_iters .................................. None
lr_decay_samples ................................ 61035
lr_decay_style .................................. linear
lr_warmup_fraction .............................. None
lr_warmup_init .................................. 0.0
lr_warmup_iters ................................. 0
lr_warmup_samples ............................... 3051
lr_wsd_decay_iters .............................. None
lr_wsd_decay_samples ............................ None
lr_wsd_decay_style .............................. exponential
main_grads_dtype ................................ torch.float32
main_params_dtype ............................... torch.float32
make_vocab_size_divisible_by .................... 128
mamba_expand .................................... 1
mamba_head_dim .................................. 64
mamba_num_groups ................................ 6
mamba_num_heads ................................. None
mamba_state_dim ................................. 320
manual_gc ....................................... False
manual_gc_eval .................................. True
manual_gc_interval .............................. 0
mask_factor ..................................... 1.0
mask_prob ....................................... 0.15
mask_type ....................................... random
masked_softmax_fusion ........................... False
max_position_embeddings ......................... 32768
max_tokens_to_oom ............................... 12000
memory_snapshot_path ............................ snapshot.pickle
merge_file ...................................... None
micro_batch_size ................................ 1
microbatch_group_size_per_vp_stage .............. None
mid_level_dataset_surplus ....................... 0.005
min_loss_scale .................................. 1.0
min_lr .......................................... 7e-07
mlp_chunks_for_prefill .......................... 1
mmap_bin_files .................................. True
mock_data ....................................... False
moe_aux_loss_coeff .............................. 0.0
moe_enable_deepep ............................... False
moe_expert_capacity_factor ...................... None
moe_extended_tp ................................. False
moe_ffn_hidden_size ............................. None
moe_grouped_gemm ................................ False
moe_input_jitter_eps ............................ None
moe_layer_freq .................................. 1
moe_layer_recompute ............................. False
moe_pad_expert_input_to_capacity ................ False
moe_per_layer_logging ........................... False
moe_permute_fusion .............................. False
moe_router_bias_update_method ................... sign
moe_router_bias_update_rate ..................... 0.001
moe_router_dtype ................................ None
moe_router_enable_expert_bias ................... False
moe_router_group_topk ........................... None
moe_router_load_balancing_type .................. aux_loss
moe_router_num_groups ........................... None
moe_router_pre_softmax .......................... False
moe_router_score_function ....................... softmax
moe_router_topk ................................. 2
moe_router_topk_scaling_factor .................. None
moe_shared_expert_intermediate_size ............. None
moe_shared_expert_overlap ....................... False
moe_token_dispatcher_type ....................... allgather
moe_token_drop_policy ........................... probs
moe_use_legacy_grouped_gemm ..................... False
moe_use_upcycling ............................... False
moe_z_loss_coeff ................................ None
mrope_section ................................... None
mscale .......................................... 1.0
mscale_all_dim .................................. 1.0
mtp_loss_scaling_factor ......................... 0.1
mtp_num_layers .................................. None
multi_latent_attention .......................... False
muon_matched_adamw_rms .......................... 0.2
muon_momentum ................................... 0.95
muon_nesterov ................................... True
muon_ns_steps ................................... 5
nccl_communicator_config_path ................... None
no_load_optim ................................... True
no_load_rng ..................................... True
no_persist_layer_norm ........................... False
no_save_optim ................................... None
no_save_rng ..................................... None
no_save_step_one ................................ True
non_persistent_ckpt_type ........................ None
non_persistent_global_ckpt_dir .................. None
non_persistent_local_ckpt_algo .................. fully_parallel
non_persistent_local_ckpt_dir ................... None
non_persistent_save_interval .................... None
norm_epsilon .................................... 1e-05
normalization ................................... RMSNorm
num_attention_heads ............................. 30
num_channels .................................... 3
num_classes ..................................... 1000
num_dataset_builder_threads ..................... 1
num_distributed_optimizer_instances ............. 1
num_experts ..................................... None
num_layers ...................................... 112
num_layers_at_end_in_bf16 ....................... 1
num_layers_at_start_in_bf16 ..................... 1
num_layers_per_virtual_pipeline_stage ........... None
num_query_groups ................................ 6
num_virtual_stages_per_pipeline_rank ............ None
num_workers ..................................... 2
object_storage_cache_path ....................... None
one_logger_async ................................ False
one_logger_project .............................. megatron-lm
one_logger_run_name ............................. None
onnx_safe ....................................... None
openai_gelu ..................................... False
optimizer ....................................... adam
optimizer_cpu_offload ........................... False
optimizer_offload_fraction ...................... 1.0
output_bert_embeddings .......................... False
overlap_cpu_optimizer_d2h_h2d ................... False
overlap_grad_reduce ............................. True
overlap_p2p_comm ................................ False
overlap_p2p_comm_warmup_flush ................... False
overlap_param_gather ............................ True
overlap_param_gather_with_optimizer_step ........ False
override_opt_param_scheduler .................... False
params_dtype .................................... torch.bfloat16
patch_dim ....................................... 16
per_split_data_args_path ........................ None
perform_initialization .......................... True
pin_cpu_grads ................................... True
pin_cpu_params .................................. True
pipeline_model_parallel_comm_backend ............ None
pipeline_model_parallel_size .................... 1
pipeline_model_parallel_split_rank .............. None
position_embedding_type ......................... rope
pretrained_checkpoint ........................... None
profile ......................................... False
profile_ranks ................................... [0]
profile_step_end ................................ 12
profile_step_start .............................. 10
q_lora_rank ..................................... None
qk_head_dim ..................................... 128
qk_l2_norm ...................................... False
qk_layernorm .................................... False
qk_pos_emb_head_dim ............................. 64
query_in_block_prob ............................. 0.1
rampup_batch_size ............................... None
rank ............................................ 0
recompute_granularity ........................... selective
recompute_method ................................ None
recompute_modules ............................... None
recompute_num_layers ............................ None
record_memory_history ........................... False
relative_attention_max_distance ................. 128
relative_attention_num_buckets .................. 32
replication ..................................... False
replication_factor .............................. 2
replication_jump ................................ None
rerun_mode ...................................... disabled
reset_attention_mask ............................ False
reset_position_ids .............................. False
result_rejected_tracker_filename ................ None
retriever_report_topk_accuracies ................ []
retriever_score_scaling ......................... False
retriever_seq_length ............................ 256
retro_add_retriever ............................. False
retro_attention_gate ............................ 1
retro_cyclic_train_iters ........................ None
retro_encoder_attention_dropout ................. 0.1
retro_encoder_hidden_dropout .................... 0.1
retro_encoder_layers ............................ 2
retro_num_neighbors ............................. 2
retro_num_retrieved_chunks ...................... 2
retro_project_dir ............................... None
retro_verify_neighbor_count ..................... True
rope_scaling_factor ............................. 8.0
rotary_base ..................................... 640000
rotary_interleaved .............................. False
rotary_percent .................................. 1.0
rotary_scaling_factor ........................... 1.0
rotary_seq_len_interpolation_factor ............. None
run_workload_inspector_server ................... False
sample_rate ..................................... 1.0
save ............................................ /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768
save_interval ................................... 29
scatter_gather_tensors_in_pipeline .............. True
seed ............................................ 1234
seq_length ...................................... 32768
sequence_parallel ............................... True
sgd_momentum .................................... 0.9
short_seq_prob .................................. 0.1
skip_data_prepare ............................... False
skip_train ...................................... False
skipped_train_samples ........................... 0
spec ............................................ ['megatron.core.models.mamba.mamba_layer_specs', 'mamba_moe_stack_spec']
split ........................................... 100,0,0
sqreglu ......................................... False
squared_relu .................................... False
start_weight_decay .............................. 0.1
straggler_ctrlr_port ............................ 65535
straggler_minmax_count .......................... 1
suggested_communication_unit_size ............... None
swiglu .......................................... True
swin_backbone_type .............................. tiny
te_rng_tracker .................................. False
tensor_model_parallel_size ...................... 2
tensorboard_dir ................................. /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/log/2025.09.17-22.24.08_based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768
tensorboard_log_interval ........................ 1
tensorboard_queue_size .......................... 1000
test_data_path .................................. None
test_mode ....................................... False
tiktoken_num_special_tokens ..................... 1000
tiktoken_pattern ................................ None
tiktoken_special_tokens ......................... None
timing_log_level ................................ 0
timing_log_option ............................... minmax
titles_data_path ................................ None
tokenizer_model ................................. /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache/models/huggingface/yulan-team/YuLan-Mini
tokenizer_type .................................. HuggingFaceTokenizer
tp_comm_bootstrap_backend ....................... nccl
tp_comm_bulk_dgrad .............................. True
tp_comm_bulk_wgrad .............................. True
tp_comm_overlap ................................. False
tp_comm_overlap_ag .............................. True
tp_comm_overlap_cfg ............................. None
tp_comm_overlap_rs .............................. True
tp_comm_overlap_rs_dgrad ........................ False
tp_comm_split_ag ................................ True
tp_comm_split_rs ................................ True
train_data_path ................................. None
train_iters ..................................... None
train_samples ................................... 61035
train_sync_interval ............................. None
transformer_impl ................................ transformer_engine
transformer_pipeline_model_parallel_size ........ 1
untie_embeddings_and_output_weights ............. True
use_checkpoint_args ............................. False
use_checkpoint_opt_param_scheduler .............. False
use_cpu_initialization .......................... None
use_custom_fsdp ................................. False
use_dist_ckpt ................................... False
use_dist_ckpt_deprecated ........................ False
use_distributed_optimizer ....................... True
use_flash_attn .................................. True
use_legacy_models ............................... False
use_mp_args_from_checkpoint_args ................ False
use_one_sent_docs ............................... False
use_persistent_ckpt_worker ...................... False
use_precision_aware_optimizer ................... False
use_pytorch_profiler ............................ False
use_ring_exchange_p2p ........................... False
use_rope_scaling ................................ False
use_rotary_position_embeddings .................. False
use_tokenizer_model_from_checkpoint_args ........ True
use_torch_fsdp2 ................................. False
use_torch_optimizer_for_cpu_offload ............. False
use_tp_pp_dp_mapping ............................ False
v_head_dim ...................................... 128
valid_data_path ................................. None
variable_seq_lengths ............................ False
virtual_pipeline_model_parallel_size ............ None
vision_backbone_type ............................ vit
vision_pretraining .............................. False
vision_pretraining_type ......................... classify
vocab_extra_ids ................................. 0
vocab_file ...................................... None
vocab_size ...................................... None
wandb_exp_name ..................................
wandb_project ...................................
wandb_save_dir ..................................
weight_decay .................................... 0.1
weight_decay_incr_style ......................... constant
wgrad_deferral_limit ............................ 0
window_size ..................................... None
world_size ...................................... 8
yaml_cfg ........................................ None
-------------------- end of arguments ---------------------
INFO:megatron.core.num_microbatches_calculator:setting number of microbatches to constant 512
> building HuggingFaceTokenizer tokenizer ...
> padded vocab (size: 99000) with 72 dummy tokens (new size: 99072)
WARNING:megatron.core.rerun_state_machine:RerunStateMachine initialized in mode RerunMode.DISABLED
> initializing torch distributed ...
> setting tensorboard ...
WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it
> initialized tensor model parallel with size 2
> initialized pipeline model parallel with size 1
> setting random seeds to 1234 ...
> compiling dataset index builder ...
[rank6]:[W917 22:25:26.615407181 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 6] using GPU 6 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device.
[rank4]:[W917 22:25:26.616202999 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 4] using GPU 4 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device.
make: Entering directory '/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/datasets'
[rank2]:[W917 22:25:26.616462451 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 2] using GPU 2 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device.
make: Nothing to be done for 'default'.
make: Leaving directory '/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/datasets'
>>> done with dataset index builder. Compilation time: 0.241 seconds
WARNING: constraints for invoking optimized fused softmax kernel are not met. We default back to unfused kernel invocations.
> compiling and loading fused kernels ...
[rank0]:[W917 22:25:27.015701688 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 0] using GPU 0 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device.
[rank5]:[W917 22:25:27.052206707 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 5] using GPU 5 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device.
[rank1]:[W917 22:25:27.054241606 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 1] using GPU 1 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device.
[rank3]:[W917 22:25:27.055513460 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 3] using GPU 3 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device.
[rank7]:[W917 22:25:27.057617657 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 7] using GPU 7 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device.
>>> done with compiling and loading fused kernels. Compilation time: 4.598 seconds
time to initialize megatron (seconds): 62.726
[after megatron is initialized] datetime: 2025-09-17 22:26:08
building Mamba model ...
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed.
warnings.warn(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed.
warnings.warn(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed.
warnings.warn(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed.
warnings.warn(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed.
warnings.warn(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed.
warnings.warn(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed.
warnings.warn(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed.
warnings.warn(
INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:Using hybrid override pattern
INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:Warning: overriding pattern A with pattern B
INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:A: M-M-M-M--M-M-*M-M-M-M-M--M-*M-M-M-M-M-M--*M-M-M-M-M-M-M-*-M-M-M-M-M-M-*M--M-M-M-M-M-*M-M--M-M-M-M-*M-M-M--M-M-M-
INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:B: *-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-
INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:Hybrid allocation (M is mamba, * is attention, - is mlp):
INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-
INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:7 attention layers in 112 total layers.
INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:Target attention ratio: 0.06. Actual attention ratio: 0.06.
INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:56 mlp layers in 112 total layers.
INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:Target mlp ratio: 0.50. Actual mlp ratio: 0.50.
- decoder.layers.0.self_attention.linear_proj.weight: 1843200
- decoder.layers.0.self_attention.linear_qkv.layer_norm_weight: 1920
- decoder.layers.0.self_attention.linear_qkv.weight: 2580480
- decoder.layers.0.self_attention.linear_qkv.bias: 1344
== params layer 0: 4426944
- decoder.layers.1.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.1.mlp.linear_fc1.weight: 9216000
- decoder.layers.1.mlp.linear_fc2.weight: 4608000
== params layer 1: 13825920
- decoder.layers.2.mixer.dt_bias: 15
- decoder.layers.2.mixer.A_log: 15
- decoder.layers.2.mixer.D: 15
- decoder.layers.2.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.2.mixer.in_proj.weight: 7401600
- decoder.layers.2.mixer.conv1d.weight: 11520
- decoder.layers.2.mixer.conv1d.bias: 2880
- decoder.layers.2.mixer.norm.weight: 960
- decoder.layers.2.mixer.out_proj.weight: 1843200
== params layer 2: 9262125
- decoder.layers.3.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.3.mlp.linear_fc1.weight: 9216000
- decoder.layers.3.mlp.linear_fc2.weight: 4608000
== params layer 3: 13825920
- decoder.layers.4.mixer.dt_bias: 15
- decoder.layers.4.mixer.A_log: 15
- decoder.layers.4.mixer.D: 15
- decoder.layers.4.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.4.mixer.in_proj.weight: 7401600
- decoder.layers.4.mixer.conv1d.weight: 11520
- decoder.layers.4.mixer.conv1d.bias: 2880
- decoder.layers.4.mixer.norm.weight: 960
- decoder.layers.4.mixer.out_proj.weight: 1843200
== params layer 4: 9262125
- decoder.layers.5.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.5.mlp.linear_fc1.weight: 9216000
- decoder.layers.5.mlp.linear_fc2.weight: 4608000
== params layer 5: 13825920
- decoder.layers.6.mixer.dt_bias: 15
- decoder.layers.6.mixer.A_log: 15
- decoder.layers.6.mixer.D: 15
- decoder.layers.6.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.6.mixer.in_proj.weight: 7401600
- decoder.layers.6.mixer.conv1d.weight: 11520
- decoder.layers.6.mixer.conv1d.bias: 2880
- decoder.layers.6.mixer.norm.weight: 960
- decoder.layers.6.mixer.out_proj.weight: 1843200
== params layer 6: 9262125
- decoder.layers.7.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.7.mlp.linear_fc1.weight: 9216000
- decoder.layers.7.mlp.linear_fc2.weight: 4608000
== params layer 7: 13825920
- decoder.layers.8.mixer.dt_bias: 15
- decoder.layers.8.mixer.A_log: 15
- decoder.layers.8.mixer.D: 15
- decoder.layers.8.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.8.mixer.in_proj.weight: 7401600
- decoder.layers.8.mixer.conv1d.weight: 11520
- decoder.layers.8.mixer.conv1d.bias: 2880
- decoder.layers.8.mixer.norm.weight: 960
- decoder.layers.8.mixer.out_proj.weight: 1843200
== params layer 8: 9262125
- decoder.layers.9.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.9.mlp.linear_fc1.weight: 9216000
- decoder.layers.9.mlp.linear_fc2.weight: 4608000
== params layer 9: 13825920
- decoder.layers.10.mixer.dt_bias: 15
- decoder.layers.10.mixer.A_log: 15
- decoder.layers.10.mixer.D: 15
- decoder.layers.10.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.10.mixer.in_proj.weight: 7401600
- decoder.layers.10.mixer.conv1d.weight: 11520
- decoder.layers.10.mixer.conv1d.bias: 2880
- decoder.layers.10.mixer.norm.weight: 960
- decoder.layers.10.mixer.out_proj.weight: 1843200
== params layer 10: 9262125
- decoder.layers.11.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.11.mlp.linear_fc1.weight: 9216000
- decoder.layers.11.mlp.linear_fc2.weight: 4608000
== params layer 11: 13825920
- decoder.layers.12.mixer.dt_bias: 15
- decoder.layers.12.mixer.A_log: 15
- decoder.layers.12.mixer.D: 15
- decoder.layers.12.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.12.mixer.in_proj.weight: 7401600
- decoder.layers.12.mixer.conv1d.weight: 11520
- decoder.layers.12.mixer.conv1d.bias: 2880
- decoder.layers.12.mixer.norm.weight: 960
- decoder.layers.12.mixer.out_proj.weight: 1843200
== params layer 12: 9262125
- decoder.layers.13.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.13.mlp.linear_fc1.weight: 9216000
- decoder.layers.13.mlp.linear_fc2.weight: 4608000
== params layer 13: 13825920
- decoder.layers.14.mixer.dt_bias: 15
- decoder.layers.14.mixer.A_log: 15
- decoder.layers.14.mixer.D: 15
- decoder.layers.14.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.14.mixer.in_proj.weight: 7401600
- decoder.layers.14.mixer.conv1d.weight: 11520
- decoder.layers.14.mixer.conv1d.bias: 2880
- decoder.layers.14.mixer.norm.weight: 960
- decoder.layers.14.mixer.out_proj.weight: 1843200
== params layer 14: 9262125
- decoder.layers.15.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.15.mlp.linear_fc1.weight: 9216000
- decoder.layers.15.mlp.linear_fc2.weight: 4608000
== params layer 15: 13825920
- decoder.layers.16.self_attention.linear_proj.weight: 1843200
- decoder.layers.16.self_attention.linear_qkv.layer_norm_weight: 1920
- decoder.layers.16.self_attention.linear_qkv.weight: 2580480
- decoder.layers.16.self_attention.linear_qkv.bias: 1344
== params layer 16: 4426944
- decoder.layers.17.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.17.mlp.linear_fc1.weight: 9216000
- decoder.layers.17.mlp.linear_fc2.weight: 4608000
== params layer 17: 13825920
- decoder.layers.18.mixer.dt_bias: 15
- decoder.layers.18.mixer.A_log: 15
- decoder.layers.18.mixer.D: 15
- decoder.layers.18.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.18.mixer.in_proj.weight: 7401600
- decoder.layers.18.mixer.conv1d.weight: 11520
- decoder.layers.18.mixer.conv1d.bias: 2880
- decoder.layers.18.mixer.norm.weight: 960
- decoder.layers.18.mixer.out_proj.weight: 1843200
== params layer 18: 9262125
- decoder.layers.19.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.19.mlp.linear_fc1.weight: 9216000
- decoder.layers.19.mlp.linear_fc2.weight: 4608000
== params layer 19: 13825920
- decoder.layers.20.mixer.dt_bias: 15
- decoder.layers.20.mixer.A_log: 15
- decoder.layers.20.mixer.D: 15
- decoder.layers.20.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.20.mixer.in_proj.weight: 7401600
- decoder.layers.20.mixer.conv1d.weight: 11520
- decoder.layers.20.mixer.conv1d.bias: 2880
- decoder.layers.20.mixer.norm.weight: 960
- decoder.layers.20.mixer.out_proj.weight: 1843200
== params layer 20: 9262125
- decoder.layers.21.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.21.mlp.linear_fc1.weight: 9216000
- decoder.layers.21.mlp.linear_fc2.weight: 4608000
== params layer 21: 13825920
- decoder.layers.22.mixer.dt_bias: 15
- decoder.layers.22.mixer.A_log: 15
- decoder.layers.22.mixer.D: 15
- decoder.layers.22.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.22.mixer.in_proj.weight: 7401600
- decoder.layers.22.mixer.conv1d.weight: 11520
- decoder.layers.22.mixer.conv1d.bias: 2880
- decoder.layers.22.mixer.norm.weight: 960
- decoder.layers.22.mixer.out_proj.weight: 1843200
== params layer 22: 9262125
- decoder.layers.23.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.23.mlp.linear_fc1.weight: 9216000
- decoder.layers.23.mlp.linear_fc2.weight: 4608000
== params layer 23: 13825920
- decoder.layers.24.mixer.dt_bias: 15
- decoder.layers.24.mixer.A_log: 15
- decoder.layers.24.mixer.D: 15
- decoder.layers.24.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.24.mixer.in_proj.weight: 7401600
- decoder.layers.24.mixer.conv1d.weight: 11520
- decoder.layers.24.mixer.conv1d.bias: 2880
- decoder.layers.24.mixer.norm.weight: 960
- decoder.layers.24.mixer.out_proj.weight: 1843200
== params layer 24: 9262125
- decoder.layers.25.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.25.mlp.linear_fc1.weight: 9216000
- decoder.layers.25.mlp.linear_fc2.weight: 4608000
== params layer 25: 13825920
- decoder.layers.26.mixer.dt_bias: 15
- decoder.layers.26.mixer.A_log: 15
- decoder.layers.26.mixer.D: 15
- decoder.layers.26.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.26.mixer.in_proj.weight: 7401600
- decoder.layers.26.mixer.conv1d.weight: 11520
- decoder.layers.26.mixer.conv1d.bias: 2880
- decoder.layers.26.mixer.norm.weight: 960
- decoder.layers.26.mixer.out_proj.weight: 1843200
== params layer 26: 9262125
- decoder.layers.27.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.27.mlp.linear_fc1.weight: 9216000
- decoder.layers.27.mlp.linear_fc2.weight: 4608000
== params layer 27: 13825920
- decoder.layers.28.mixer.dt_bias: 15
- decoder.layers.28.mixer.A_log: 15
- decoder.layers.28.mixer.D: 15
- decoder.layers.28.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.28.mixer.in_proj.weight: 7401600
- decoder.layers.28.mixer.conv1d.weight: 11520
- decoder.layers.28.mixer.conv1d.bias: 2880
- decoder.layers.28.mixer.norm.weight: 960
- decoder.layers.28.mixer.out_proj.weight: 1843200
== params layer 28: 9262125
- decoder.layers.29.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.29.mlp.linear_fc1.weight: 9216000
- decoder.layers.29.mlp.linear_fc2.weight: 4608000
== params layer 29: 13825920
- decoder.layers.30.mixer.dt_bias: 15
- decoder.layers.30.mixer.A_log: 15
- decoder.layers.30.mixer.D: 15
- decoder.layers.30.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.30.mixer.in_proj.weight: 7401600
- decoder.layers.30.mixer.conv1d.weight: 11520
- decoder.layers.30.mixer.conv1d.bias: 2880
- decoder.layers.30.mixer.norm.weight: 960
- decoder.layers.30.mixer.out_proj.weight: 1843200
== params layer 30: 9262125
- decoder.layers.31.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.31.mlp.linear_fc1.weight: 9216000
- decoder.layers.31.mlp.linear_fc2.weight: 4608000
== params layer 31: 13825920
- decoder.layers.32.self_attention.linear_proj.weight: 1843200
- decoder.layers.32.self_attention.linear_qkv.layer_norm_weight: 1920
- decoder.layers.32.self_attention.linear_qkv.weight: 2580480
- decoder.layers.32.self_attention.linear_qkv.bias: 1344
== params layer 32: 4426944
- decoder.layers.33.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.33.mlp.linear_fc1.weight: 9216000
- decoder.layers.33.mlp.linear_fc2.weight: 4608000
== params layer 33: 13825920
- decoder.layers.34.mixer.dt_bias: 15
- decoder.layers.34.mixer.A_log: 15
- decoder.layers.34.mixer.D: 15
- decoder.layers.34.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.34.mixer.in_proj.weight: 7401600
- decoder.layers.34.mixer.conv1d.weight: 11520
- decoder.layers.34.mixer.conv1d.bias: 2880
- decoder.layers.34.mixer.norm.weight: 960
- decoder.layers.34.mixer.out_proj.weight: 1843200
== params layer 34: 9262125
- decoder.layers.35.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.35.mlp.linear_fc1.weight: 9216000
- decoder.layers.35.mlp.linear_fc2.weight: 4608000
== params layer 35: 13825920
- decoder.layers.36.mixer.dt_bias: 15
- decoder.layers.36.mixer.A_log: 15
- decoder.layers.36.mixer.D: 15
- decoder.layers.36.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.36.mixer.in_proj.weight: 7401600
- decoder.layers.36.mixer.conv1d.weight: 11520
- decoder.layers.36.mixer.conv1d.bias: 2880
- decoder.layers.36.mixer.norm.weight: 960
- decoder.layers.36.mixer.out_proj.weight: 1843200
== params layer 36: 9262125
- decoder.layers.37.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.37.mlp.linear_fc1.weight: 9216000
- decoder.layers.37.mlp.linear_fc2.weight: 4608000
== params layer 37: 13825920
- decoder.layers.38.mixer.dt_bias: 15
- decoder.layers.38.mixer.A_log: 15
- decoder.layers.38.mixer.D: 15
- decoder.layers.38.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.38.mixer.in_proj.weight: 7401600
- decoder.layers.38.mixer.conv1d.weight: 11520
- decoder.layers.38.mixer.conv1d.bias: 2880
- decoder.layers.38.mixer.norm.weight: 960
- decoder.layers.38.mixer.out_proj.weight: 1843200
== params layer 38: 9262125
- decoder.layers.39.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.39.mlp.linear_fc1.weight: 9216000
- decoder.layers.39.mlp.linear_fc2.weight: 4608000
== params layer 39: 13825920
- decoder.layers.40.mixer.dt_bias: 15
- decoder.layers.40.mixer.A_log: 15
- decoder.layers.40.mixer.D: 15
- decoder.layers.40.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.40.mixer.in_proj.weight: 7401600
- decoder.layers.40.mixer.conv1d.weight: 11520
- decoder.layers.40.mixer.conv1d.bias: 2880
- decoder.layers.40.mixer.norm.weight: 960
- decoder.layers.40.mixer.out_proj.weight: 1843200
== params layer 40: 9262125
- decoder.layers.41.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.41.mlp.linear_fc1.weight: 9216000
- decoder.layers.41.mlp.linear_fc2.weight: 4608000
== params layer 41: 13825920
- decoder.layers.42.mixer.dt_bias: 15
- decoder.layers.42.mixer.A_log: 15
- decoder.layers.42.mixer.D: 15
- decoder.layers.42.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.42.mixer.in_proj.weight: 7401600
- decoder.layers.42.mixer.conv1d.weight: 11520
- decoder.layers.42.mixer.conv1d.bias: 2880
- decoder.layers.42.mixer.norm.weight: 960
- decoder.layers.42.mixer.out_proj.weight: 1843200
== params layer 42: 9262125
- decoder.layers.43.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.43.mlp.linear_fc1.weight: 9216000
- decoder.layers.43.mlp.linear_fc2.weight: 4608000
> number of parameters on (tensor, pipeline) model parallel rank (1, 0): 1449304413
== params layer 43: 13825920
- decoder.layers.44.mixer.dt_bias: 15
- decoder.layers.44.mixer.A_log: 15
- decoder.layers.44.mixer.D: 15
- decoder.layers.44.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.44.mixer.in_proj.weight: 7401600
- decoder.layers.44.mixer.conv1d.weight: 11520
- decoder.layers.44.mixer.conv1d.bias: 2880
- decoder.layers.44.mixer.norm.weight: 960
- decoder.layers.44.mixer.out_proj.weight: 1843200
== params layer 44: 9262125
- decoder.layers.45.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.45.mlp.linear_fc1.weight: 9216000
- decoder.layers.45.mlp.linear_fc2.weight: 4608000
== params layer 45: 13825920
- decoder.layers.46.mixer.dt_bias: 15
- decoder.layers.46.mixer.A_log: 15
- decoder.layers.46.mixer.D: 15
- decoder.layers.46.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.46.mixer.in_proj.weight: 7401600
- decoder.layers.46.mixer.conv1d.weight: 11520
- decoder.layers.46.mixer.conv1d.bias: 2880
- decoder.layers.46.mixer.norm.weight: 960
- decoder.layers.46.mixer.out_proj.weight: 1843200
== params layer 46: 9262125
- decoder.layers.47.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.47.mlp.linear_fc1.weight: 9216000
- decoder.layers.47.mlp.linear_fc2.weight: 4608000
== params layer 47: 13825920
- decoder.layers.48.self_attention.linear_proj.weight: 1843200
- decoder.layers.48.self_attention.linear_qkv.layer_norm_weight: 1920
- decoder.layers.48.self_attention.linear_qkv.weight: 2580480
- decoder.layers.48.self_attention.linear_qkv.bias: 1344
== params layer 48: 4426944
- decoder.layers.49.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.49.mlp.linear_fc1.weight: 9216000
- decoder.layers.49.mlp.linear_fc2.weight: 4608000
== params layer 49: 13825920
- decoder.layers.50.mixer.dt_bias: 15
- decoder.layers.50.mixer.A_log: 15
- decoder.layers.50.mixer.D: 15
- decoder.layers.50.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.50.mixer.in_proj.weight: 7401600
- decoder.layers.50.mixer.conv1d.weight: 11520
- decoder.layers.50.mixer.conv1d.bias: 2880
- decoder.layers.50.mixer.norm.weight: 960
- decoder.layers.50.mixer.out_proj.weight: 1843200
== params layer 50: 9262125
- decoder.layers.51.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.51.mlp.linear_fc1.weight: 9216000
- decoder.layers.51.mlp.linear_fc2.weight: 4608000
== params layer 51: 13825920
- decoder.layers.52.mixer.dt_bias: 15
- decoder.layers.52.mixer.A_log: 15
- decoder.layers.52.mixer.D: 15
- decoder.layers.52.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.52.mixer.in_proj.weight: 7401600
- decoder.layers.52.mixer.conv1d.weight: 11520
- decoder.layers.52.mixer.conv1d.bias: 2880
- decoder.layers.52.mixer.norm.weight: 960
- decoder.layers.52.mixer.out_proj.weight: 1843200
== params layer 52: 9262125
- decoder.layers.53.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.53.mlp.linear_fc1.weight: 9216000
- decoder.layers.53.mlp.linear_fc2.weight: 4608000
== params layer 53: 13825920
- decoder.layers.54.mixer.dt_bias: 15
- decoder.layers.54.mixer.A_log: 15
- decoder.layers.54.mixer.D: 15
- decoder.layers.54.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.54.mixer.in_proj.weight: 7401600
- decoder.layers.54.mixer.conv1d.weight: 11520
- decoder.layers.54.mixer.conv1d.bias: 2880
- decoder.layers.54.mixer.norm.weight: 960
- decoder.layers.54.mixer.out_proj.weight: 1843200
== params layer 54: 9262125
- decoder.layers.55.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.55.mlp.linear_fc1.weight: 9216000
- decoder.layers.55.mlp.linear_fc2.weight: 4608000
== params layer 55: 13825920
- decoder.layers.56.mixer.dt_bias: 15
- decoder.layers.56.mixer.A_log: 15
- decoder.layers.56.mixer.D: 15
- decoder.layers.56.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.56.mixer.in_proj.weight: 7401600
- decoder.layers.56.mixer.conv1d.weight: 11520
- decoder.layers.56.mixer.conv1d.bias: 2880
- decoder.layers.56.mixer.norm.weight: 960
- decoder.layers.56.mixer.out_proj.weight: 1843200
== params layer 56: 9262125
- decoder.layers.57.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.57.mlp.linear_fc1.weight: 9216000
- decoder.layers.57.mlp.linear_fc2.weight: 4608000
== params layer 57: 13825920
- decoder.layers.58.mixer.dt_bias: 15
- decoder.layers.58.mixer.A_log: 15
- decoder.layers.58.mixer.D: 15
- decoder.layers.58.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.58.mixer.in_proj.weight: 7401600
- decoder.layers.58.mixer.conv1d.weight: 11520
- decoder.layers.58.mixer.conv1d.bias: 2880
- decoder.layers.58.mixer.norm.weight: 960
- decoder.layers.58.mixer.out_proj.weight: 1843200
== params layer 58: 9262125
- decoder.layers.59.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.59.mlp.linear_fc1.weight: 9216000
- decoder.layers.59.mlp.linear_fc2.weight: 4608000
== params layer 59: 13825920
- decoder.layers.60.mixer.dt_bias: 15
- decoder.layers.60.mixer.A_log: 15
- decoder.layers.60.mixer.D: 15
- decoder.layers.60.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.60.mixer.in_proj.weight: 7401600
- decoder.layers.60.mixer.conv1d.weight: 11520
- decoder.layers.60.mixer.conv1d.bias: 2880
- decoder.layers.60.mixer.norm.weight: 960
- decoder.layers.60.mixer.out_proj.weight: 1843200
== params layer 60: 9262125
- decoder.layers.61.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.61.mlp.linear_fc1.weight: 9216000
- decoder.layers.61.mlp.linear_fc2.weight: 4608000
== params layer 61: 13825920
- decoder.layers.62.mixer.dt_bias: 15
- decoder.layers.62.mixer.A_log: 15
- decoder.layers.62.mixer.D: 15
- decoder.layers.62.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.62.mixer.in_proj.weight: 7401600
- decoder.layers.62.mixer.conv1d.weight: 11520
- decoder.layers.62.mixer.conv1d.bias: 2880
- decoder.layers.62.mixer.norm.weight: 960
- decoder.layers.62.mixer.out_proj.weight: 1843200
== params layer 62: 9262125
- decoder.layers.63.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.63.mlp.linear_fc1.weight: 9216000
- decoder.layers.63.mlp.linear_fc2.weight: 4608000
== params layer 63: 13825920
- decoder.layers.64.self_attention.linear_proj.weight: 1843200
- decoder.layers.64.self_attention.linear_qkv.layer_norm_weight: 1920
- decoder.layers.64.self_attention.linear_qkv.weight: 2580480
- decoder.layers.64.self_attention.linear_qkv.bias: 1344
== params layer 64: 4426944
- decoder.layers.65.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.65.mlp.linear_fc1.weight: 9216000
- decoder.layers.65.mlp.linear_fc2.weight: 4608000
== params layer 65: 13825920
- decoder.layers.66.mixer.dt_bias: 15
- decoder.layers.66.mixer.A_log: 15
- decoder.layers.66.mixer.D: 15
- decoder.layers.66.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.66.mixer.in_proj.weight: 7401600
- decoder.layers.66.mixer.conv1d.weight: 11520
- decoder.layers.66.mixer.conv1d.bias: 2880
- decoder.layers.66.mixer.norm.weight: 960
- decoder.layers.66.mixer.out_proj.weight: 1843200
== params layer 66: 9262125
- decoder.layers.67.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.67.mlp.linear_fc1.weight: 9216000
- decoder.layers.67.mlp.linear_fc2.weight: 4608000
== params layer 67: 13825920
- decoder.layers.68.mixer.dt_bias: 15
- decoder.layers.68.mixer.A_log: 15
- decoder.layers.68.mixer.D: 15
- decoder.layers.68.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.68.mixer.in_proj.weight: 7401600
- decoder.layers.68.mixer.conv1d.weight: 11520
- decoder.layers.68.mixer.conv1d.bias: 2880
- decoder.layers.68.mixer.norm.weight: 960
- decoder.layers.68.mixer.out_proj.weight: 1843200
== params layer 68: 9262125
- decoder.layers.69.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.69.mlp.linear_fc1.weight: 9216000
- decoder.layers.69.mlp.linear_fc2.weight: 4608000
== params layer 69: 13825920
- decoder.layers.70.mixer.dt_bias: 15
- decoder.layers.70.mixer.A_log: 15
- decoder.layers.70.mixer.D: 15
- decoder.layers.70.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.70.mixer.in_proj.weight: 7401600
- decoder.layers.70.mixer.conv1d.weight: 11520
- decoder.layers.70.mixer.conv1d.bias: 2880
- decoder.layers.70.mixer.norm.weight: 960
- decoder.layers.70.mixer.out_proj.weight: 1843200
== params layer 70: 9262125
- decoder.layers.71.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.71.mlp.linear_fc1.weight: 9216000
- decoder.layers.71.mlp.linear_fc2.weight: 4608000
== params layer 71: 13825920
- decoder.layers.72.mixer.dt_bias: 15
- decoder.layers.72.mixer.A_log: 15
- decoder.layers.72.mixer.D: 15
- decoder.layers.72.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.72.mixer.in_proj.weight: 7401600
- decoder.layers.72.mixer.conv1d.weight: 11520
- decoder.layers.72.mixer.conv1d.bias: 2880
- decoder.layers.72.mixer.norm.weight: 960
- decoder.layers.72.mixer.out_proj.weight: 1843200
== params layer 72: 9262125
- decoder.layers.73.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.73.mlp.linear_fc1.weight: 9216000
- decoder.layers.73.mlp.linear_fc2.weight: 4608000
== params layer 73: 13825920
- decoder.layers.74.mixer.dt_bias: 15
- decoder.layers.74.mixer.A_log: 15
- decoder.layers.74.mixer.D: 15
- decoder.layers.74.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.74.mixer.in_proj.weight: 7401600
- decoder.layers.74.mixer.conv1d.weight: 11520
- decoder.layers.74.mixer.conv1d.bias: 2880
- decoder.layers.74.mixer.norm.weight: 960
- decoder.layers.74.mixer.out_proj.weight: 1843200
== params layer 74: 9262125
- decoder.layers.75.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.75.mlp.linear_fc1.weight: 9216000
- decoder.layers.75.mlp.linear_fc2.weight: 4608000
== params layer 75: 13825920
- decoder.layers.76.mixer.dt_bias: 15
- decoder.layers.76.mixer.A_log: 15
- decoder.layers.76.mixer.D: 15
- decoder.layers.76.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.76.mixer.in_proj.weight: 7401600
- decoder.layers.76.mixer.conv1d.weight: 11520
- decoder.layers.76.mixer.conv1d.bias: 2880
- decoder.layers.76.mixer.norm.weight: 960
- decoder.layers.76.mixer.out_proj.weight: 1843200
== params layer 76: 9262125
- decoder.layers.77.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.77.mlp.linear_fc1.weight: 9216000
- decoder.layers.77.mlp.linear_fc2.weight: 4608000
== params layer 77: 13825920
- decoder.layers.78.mixer.dt_bias: 15
- decoder.layers.78.mixer.A_log: 15
- decoder.layers.78.mixer.D: 15
- decoder.layers.78.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.78.mixer.in_proj.weight: 7401600
- decoder.layers.78.mixer.conv1d.weight: 11520
- decoder.layers.78.mixer.conv1d.bias: 2880
- decoder.layers.78.mixer.norm.weight: 960
- decoder.layers.78.mixer.out_proj.weight: 1843200
== params layer 78: 9262125
- decoder.layers.79.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.79.mlp.linear_fc1.weight: 9216000
- decoder.layers.79.mlp.linear_fc2.weight: 4608000
== params layer 79: 13825920
- decoder.layers.80.self_attention.linear_proj.weight: 1843200
- decoder.layers.80.self_attention.linear_qkv.layer_norm_weight: 1920
- decoder.layers.80.self_attention.linear_qkv.weight: 2580480
- decoder.layers.80.self_attention.linear_qkv.bias: 1344
== params layer 80: 4426944
- decoder.layers.81.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.81.mlp.linear_fc1.weight: 9216000
- decoder.layers.81.mlp.linear_fc2.weight: 4608000
== params layer 81: 13825920
- decoder.layers.82.mixer.dt_bias: 15
- decoder.layers.82.mixer.A_log: 15
- decoder.layers.82.mixer.D: 15
- decoder.layers.82.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.82.mixer.in_proj.weight: 7401600
- decoder.layers.82.mixer.conv1d.weight: 11520
- decoder.layers.82.mixer.conv1d.bias: 2880
- decoder.layers.82.mixer.norm.weight: 960
- decoder.layers.82.mixer.out_proj.weight: 1843200
== params layer 82: 9262125
- decoder.layers.83.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.83.mlp.linear_fc1.weight: 9216000
- decoder.layers.83.mlp.linear_fc2.weight: 4608000
== params layer 83: 13825920
- decoder.layers.84.mixer.dt_bias: 15
- decoder.layers.84.mixer.A_log: 15
- decoder.layers.84.mixer.D: 15
- decoder.layers.84.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.84.mixer.in_proj.weight: 7401600
- decoder.layers.84.mixer.conv1d.weight: 11520
- decoder.layers.84.mixer.conv1d.bias: 2880
- decoder.layers.84.mixer.norm.weight: 960
- decoder.layers.84.mixer.out_proj.weight: 1843200
== params layer 84: 9262125
- decoder.layers.85.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.85.mlp.linear_fc1.weight: 9216000
- decoder.layers.85.mlp.linear_fc2.weight: 4608000
== params layer 85: 13825920
- decoder.layers.86.mixer.dt_bias: 15
- decoder.layers.86.mixer.A_log: 15
- decoder.layers.86.mixer.D: 15
- decoder.layers.86.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.86.mixer.in_proj.weight: 7401600
- decoder.layers.86.mixer.conv1d.weight: 11520
- decoder.layers.86.mixer.conv1d.bias: 2880
- decoder.layers.86.mixer.norm.weight: 960
- decoder.layers.86.mixer.out_proj.weight: 1843200
== params layer 86: 9262125
- decoder.layers.87.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.87.mlp.linear_fc1.weight: 9216000
- decoder.layers.87.mlp.linear_fc2.weight: 4608000
== params layer 87: 13825920
- decoder.layers.88.mixer.dt_bias: 15
- decoder.layers.88.mixer.A_log: 15
- decoder.layers.88.mixer.D: 15
- decoder.layers.88.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.88.mixer.in_proj.weight: 7401600
- decoder.layers.88.mixer.conv1d.weight: 11520
- decoder.layers.88.mixer.conv1d.bias: 2880
- decoder.layers.88.mixer.norm.weight: 960
- decoder.layers.88.mixer.out_proj.weight: 1843200
== params layer 88: 9262125
- decoder.layers.89.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.89.mlp.linear_fc1.weight: 9216000
- decoder.layers.89.mlp.linear_fc2.weight: 4608000
== params layer 89: 13825920
- decoder.layers.90.mixer.dt_bias: 15
- decoder.layers.90.mixer.A_log: 15
- decoder.layers.90.mixer.D: 15
- decoder.layers.90.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.90.mixer.in_proj.weight: 7401600
- decoder.layers.90.mixer.conv1d.weight: 11520
- decoder.layers.90.mixer.conv1d.bias: 2880
- decoder.layers.90.mixer.norm.weight: 960
- decoder.layers.90.mixer.out_proj.weight: 1843200
== params layer 90: 9262125
- decoder.layers.91.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.91.mlp.linear_fc1.weight: 9216000
- decoder.layers.91.mlp.linear_fc2.weight: 4608000
== params layer 91: 13825920
- decoder.layers.92.mixer.dt_bias: 15
- decoder.layers.92.mixer.A_log: 15
- decoder.layers.92.mixer.D: 15
- decoder.layers.92.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.92.mixer.in_proj.weight: 7401600
- decoder.layers.92.mixer.conv1d.weight: 11520
- decoder.layers.92.mixer.conv1d.bias: 2880
- decoder.layers.92.mixer.norm.weight: 960
- decoder.layers.92.mixer.out_proj.weight: 1843200
== params layer 92: 9262125
- decoder.layers.93.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.93.mlp.linear_fc1.weight: 9216000
- decoder.layers.93.mlp.linear_fc2.weight: 4608000
== params layer 93: 13825920
- decoder.layers.94.mixer.dt_bias: 15
- decoder.layers.94.mixer.A_log: 15
- decoder.layers.94.mixer.D: 15
- decoder.layers.94.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.94.mixer.in_proj.weight: 7401600
- decoder.layers.94.mixer.conv1d.weight: 11520
- decoder.layers.94.mixer.conv1d.bias: 2880
- decoder.layers.94.mixer.norm.weight: 960
- decoder.layers.94.mixer.out_proj.weight: 1843200
== params layer 94: 9262125
- decoder.layers.95.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.95.mlp.linear_fc1.weight: 9216000
- decoder.layers.95.mlp.linear_fc2.weight: 4608000
== params layer 95: 13825920
- decoder.layers.96.self_attention.linear_proj.weight: 1843200
- decoder.layers.96.self_attention.linear_qkv.layer_norm_weight: 1920
- decoder.layers.96.self_attention.linear_qkv.weight: 2580480
- decoder.layers.96.self_attention.linear_qkv.bias: 1344
== params layer 96: 4426944
- decoder.layers.97.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.97.mlp.linear_fc1.weight: 9216000
- decoder.layers.97.mlp.linear_fc2.weight: 4608000
== params layer 97: 13825920
- decoder.layers.98.mixer.dt_bias: 15
- decoder.layers.98.mixer.A_log: 15
- decoder.layers.98.mixer.D: 15
- decoder.layers.98.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.98.mixer.in_proj.weight: 7401600
- decoder.layers.98.mixer.conv1d.weight: 11520
- decoder.layers.98.mixer.conv1d.bias: 2880
- decoder.layers.98.mixer.norm.weight: 960
- decoder.layers.98.mixer.out_proj.weight: 1843200
== params layer 98: 9262125
- decoder.layers.99.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.99.mlp.linear_fc1.weight: 9216000
- decoder.layers.99.mlp.linear_fc2.weight: 4608000
== params layer 99: 13825920
- decoder.layers.100.mixer.dt_bias: 15
- decoder.layers.100.mixer.A_log: 15
- decoder.layers.100.mixer.D: 15
- decoder.layers.100.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.100.mixer.in_proj.weight: 7401600
- decoder.layers.100.mixer.conv1d.weight: 11520
- decoder.layers.100.mixer.conv1d.bias: 2880
- decoder.layers.100.mixer.norm.weight: 960
- decoder.layers.100.mixer.out_proj.weight: 1843200
== params layer 100: 9262125
- decoder.layers.101.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.101.mlp.linear_fc1.weight: 9216000
- decoder.layers.101.mlp.linear_fc2.weight: 4608000
== params layer 101: 13825920
- decoder.layers.102.mixer.dt_bias: 15
- decoder.layers.102.mixer.A_log: 15
- decoder.layers.102.mixer.D: 15
- decoder.layers.102.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.102.mixer.in_proj.weight: 7401600
- decoder.layers.102.mixer.conv1d.weight: 11520
- decoder.layers.102.mixer.conv1d.bias: 2880
- decoder.layers.102.mixer.norm.weight: 960
- decoder.layers.102.mixer.out_proj.weight: 1843200
== params layer 102: 9262125
- decoder.layers.103.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.103.mlp.linear_fc1.weight: 9216000
- decoder.layers.103.mlp.linear_fc2.weight: 4608000
== params layer 103: 13825920
- decoder.layers.104.mixer.dt_bias: 15
- decoder.layers.104.mixer.A_log: 15
- decoder.layers.104.mixer.D: 15
- decoder.layers.104.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.104.mixer.in_proj.weight: 7401600
- decoder.layers.104.mixer.conv1d.weight: 11520
- decoder.layers.104.mixer.conv1d.bias: 2880
- decoder.layers.104.mixer.norm.weight: 960
- decoder.layers.104.mixer.out_proj.weight: 1843200
== params layer 104: 9262125
- decoder.layers.105.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.105.mlp.linear_fc1.weight: 9216000
- decoder.layers.105.mlp.linear_fc2.weight: 4608000
== params layer 105: 13825920
- decoder.layers.106.mixer.dt_bias: 15
- decoder.layers.106.mixer.A_log: 15
- decoder.layers.106.mixer.D: 15
- decoder.layers.106.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.106.mixer.in_proj.weight: 7401600
- decoder.layers.106.mixer.conv1d.weight: 11520
- decoder.layers.106.mixer.conv1d.bias: 2880
- decoder.layers.106.mixer.norm.weight: 960
- decoder.layers.106.mixer.out_proj.weight: 1843200
== params layer 106: 9262125
- decoder.layers.107.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.107.mlp.linear_fc1.weight: 9216000
- decoder.layers.107.mlp.linear_fc2.weight: 4608000
== params layer 107: 13825920
- decoder.layers.108.mixer.dt_bias: 15
- decoder.layers.108.mixer.A_log: 15
- decoder.layers.108.mixer.D: 15
- decoder.layers.108.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.108.mixer.in_proj.weight: 7401600
- decoder.layers.108.mixer.conv1d.weight: 11520
- decoder.layers.108.mixer.conv1d.bias: 2880
- decoder.layers.108.mixer.norm.weight: 960
- decoder.layers.108.mixer.out_proj.weight: 1843200
== params layer 108: 9262125
- decoder.layers.109.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.109.mlp.linear_fc1.weight: 9216000
- decoder.layers.109.mlp.linear_fc2.weight: 4608000
== params layer 109: 13825920
- decoder.layers.110.mixer.dt_bias: 15
- decoder.layers.110.mixer.A_log: 15
- decoder.layers.110.mixer.D: 15
- decoder.layers.110.mixer.in_proj.layer_norm_weight: 1920
- decoder.layers.110.mixer.in_proj.weight: 7401600
- decoder.layers.110.mixer.conv1d.weight: 11520
- decoder.layers.110.mixer.conv1d.bias: 2880
- decoder.layers.110.mixer.norm.weight: 960
- decoder.layers.110.mixer.out_proj.weight: 1843200
== params layer 110: 9262125
- decoder.layers.111.mlp.linear_fc1.layer_norm_weight: 1920
- decoder.layers.111.mlp.linear_fc1.weight: 9216000
- decoder.layers.111.mlp.linear_fc2.weight: 4608000
== params layer 111: 13825920
> number of parameters on (tensor, pipeline) model parallel rank (0, 0): 1449304413
[DEBUG] freeze_non_mamba: False
[DEBUG] freeze_non_mamba: False
> number of parameters on (tensor, pipeline) model parallel rank (1, 0): 1449304413
INFO:megatron.core.distributed.distributed_data_parallel:Setting up DistributedDataParallel with config DistributedDataParallelConfig(grad_reduce_in_fp32=True, overlap_grad_reduce=True, overlap_param_gather=True, align_param_gather=False, use_distributed_optimizer=True, num_distributed_optimizer_instances=1, check_for_nan_in_grad=True, check_for_large_grads=False, bucket_size=40000000, pad_buckets_for_high_nccl_busbw=False, average_in_collective=False, fp8_param_gather=False, use_custom_fsdp=False, data_parallel_sharding_strategy='no_shard', gradient_reduce_div_fusion=True, suggested_communication_unit_size=None, preserve_fp32_weights=True, keep_fp8_transpose_cache_when_using_custom_fsdp=False)
> number of parameters on (tensor, pipeline) model parallel rank (0, 0): 1449304413
[DEBUG] freeze_non_mamba: False
INFO:megatron.core.distributed.param_and_grad_buffer:Number of buckets for gradient all-reduce / reduce-scatter: 30
Params for bucket 1 (95109120 elements, 95109120 padded size):
module.output_layer.weight
Params for bucket 2 (46176045 elements, 46176256 padded size):
module.decoder.layers.109.mlp.linear_fc1.weight
module.decoder.layers.108.mixer.out_proj.weight
module.decoder.layers.111.mlp.linear_fc1.weight
module.decoder.layers.110.mixer.out_proj.weight
module.decoder.layers.110.mixer.conv1d.weight
module.decoder.layers.110.mixer.in_proj.layer_norm_weight
module.decoder.layers.108.mixer.conv1d.weight
module.decoder.layers.110.mixer.dt_bias
module.decoder.layers.108.mixer.conv1d.bias
module.decoder.layers.110.mixer.conv1d.bias
module.decoder.layers.109.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.110.mixer.D
module.decoder.layers.110.mixer.A_log
module.decoder.layers.111.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.108.mixer.norm.weight
module.decoder.final_norm.weight
module.decoder.layers.111.mlp.linear_fc2.weight
module.decoder.layers.110.mixer.norm.weight
module.decoder.layers.110.mixer.in_proj.weight
module.decoder.layers.109.mlp.linear_fc2.weight
module.decoder.layers.108.mixer.in_proj.weight
Params for bucket 3 (46176090 elements, 46176384 padded size):
module.decoder.layers.106.mixer.in_proj.layer_norm_weight
module.decoder.layers.106.mixer.D
module.decoder.layers.105.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.108.mixer.in_proj.layer_norm_weight
module.decoder.layers.107.mlp.linear_fc1.weight
module.decoder.layers.106.mixer.out_proj.weight
module.decoder.layers.106.mixer.conv1d.bias
module.decoder.layers.106.mixer.dt_bias
module.decoder.layers.105.mlp.linear_fc2.weight
module.decoder.layers.108.mixer.dt_bias
module.decoder.layers.104.mixer.in_proj.weight
module.decoder.layers.106.mixer.A_log
module.decoder.layers.108.mixer.A_log
module.decoder.layers.108.mixer.D
module.decoder.layers.107.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.106.mixer.norm.weight
module.decoder.layers.105.mlp.linear_fc1.weight
module.decoder.layers.104.mixer.out_proj.weight
module.decoder.layers.104.mixer.norm.weight
module.decoder.layers.104.mixer.conv1d.weight
module.decoder.layers.106.mixer.in_proj.weight
module.decoder.layers.107.mlp.linear_fc2.weight
module.decoder.layers.106.mixer.conv1d.weight
module.decoder.layers.104.mixer.conv1d.bias
Params for bucket 4 (46176090 elements, 46176384 padded size):
module.decoder.layers.102.mixer.in_proj.layer_norm_weight
module.decoder.layers.102.mixer.D
module.decoder.layers.102.mixer.A_log
module.decoder.layers.101.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.104.mixer.in_proj.layer_norm_weight
module.decoder.layers.104.mixer.A_log
module.decoder.layers.104.mixer.D
module.decoder.layers.103.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.102.mixer.in_proj.weight
module.decoder.layers.101.mlp.linear_fc2.weight
module.decoder.layers.103.mlp.linear_fc2.weight
module.decoder.layers.100.mixer.in_proj.weight
module.decoder.layers.103.mlp.linear_fc1.weight
module.decoder.layers.102.mixer.out_proj.weight
module.decoder.layers.102.mixer.norm.weight
module.decoder.layers.102.mixer.conv1d.weight
module.decoder.layers.101.mlp.linear_fc1.weight
module.decoder.layers.100.mixer.out_proj.weight
module.decoder.layers.100.mixer.norm.weight
module.decoder.layers.100.mixer.conv1d.weight
module.decoder.layers.102.mixer.dt_bias
module.decoder.layers.104.mixer.dt_bias
module.decoder.layers.102.mixer.conv1d.bias
module.decoder.layers.100.mixer.conv1d.bias
Params for bucket 5 (41342874 elements, 41343232 padded size):
module.decoder.layers.98.mixer.A_log
module.decoder.layers.100.mixer.in_proj.layer_norm_weight
module.decoder.layers.100.mixer.A_log
module.decoder.layers.100.mixer.D
module.decoder.layers.99.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.98.mixer.norm.weight
module.decoder.layers.98.mixer.D
module.decoder.layers.96.self_attention.linear_qkv.weight
module.decoder.layers.99.mlp.linear_fc2.weight
module.decoder.layers.98.mixer.conv1d.bias
module.decoder.layers.98.mixer.in_proj.weight
module.decoder.layers.97.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.96.self_attention.linear_qkv.bias
module.decoder.layers.96.self_attention.linear_proj.weight
module.decoder.layers.97.mlp.linear_fc1.weight
module.decoder.layers.98.mixer.conv1d.weight
module.decoder.layers.99.mlp.linear_fc1.weight
module.decoder.layers.98.mixer.out_proj.weight
module.decoder.layers.98.mixer.in_proj.layer_norm_weight
module.decoder.layers.98.mixer.dt_bias
module.decoder.layers.97.mlp.linear_fc2.weight
module.decoder.layers.100.mixer.dt_bias
module.decoder.layers.96.self_attention.linear_qkv.layer_norm_weight
Params for bucket 6 (46174125 elements, 46174336 padded size):
module.decoder.layers.95.mlp.linear_fc2.weight
module.decoder.layers.94.mixer.norm.weight
module.decoder.layers.94.mixer.in_proj.weight
module.decoder.layers.92.mixer.norm.weight
module.decoder.layers.93.mlp.linear_fc2.weight
module.decoder.layers.92.mixer.in_proj.weight
module.decoder.layers.93.mlp.linear_fc1.weight
module.decoder.layers.94.mixer.out_proj.weight
module.decoder.layers.95.mlp.linear_fc1.weight
module.decoder.layers.94.mixer.conv1d.weight
module.decoder.layers.94.mixer.in_proj.layer_norm_weight
module.decoder.layers.92.mixer.conv1d.weight
module.decoder.layers.92.mixer.out_proj.weight
module.decoder.layers.92.mixer.conv1d.bias
module.decoder.layers.94.mixer.dt_bias
module.decoder.layers.94.mixer.conv1d.bias
module.decoder.layers.93.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.95.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.94.mixer.D
module.decoder.layers.94.mixer.A_log
Params for bucket 7 (46176090 elements, 46176384 padded size):
module.decoder.layers.90.mixer.in_proj.weight
module.decoder.layers.91.mlp.linear_fc2.weight
module.decoder.layers.90.mixer.norm.weight
module.decoder.layers.89.mlp.linear_fc2.weight
module.decoder.layers.88.mixer.norm.weight
module.decoder.layers.88.mixer.in_proj.weight
module.decoder.layers.90.mixer.in_proj.layer_norm_weight
module.decoder.layers.90.mixer.conv1d.weight
module.decoder.layers.92.mixer.in_proj.layer_norm_weight
module.decoder.layers.91.mlp.linear_fc1.weight
module.decoder.layers.90.mixer.out_proj.weight
module.decoder.layers.89.mlp.linear_fc1.weight
module.decoder.layers.88.mixer.out_proj.weight
module.decoder.layers.88.mixer.conv1d.weight
module.decoder.layers.90.mixer.dt_bias
module.decoder.layers.92.mixer.dt_bias
module.decoder.layers.90.mixer.conv1d.bias
module.decoder.layers.88.mixer.conv1d.bias
module.decoder.layers.89.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.90.mixer.A_log
module.decoder.layers.92.mixer.D
module.decoder.layers.92.mixer.A_log
module.decoder.layers.91.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.90.mixer.D
Params for bucket 8 (46176090 elements, 46176384 padded size):
module.decoder.layers.86.mixer.in_proj.weight
module.decoder.layers.87.mlp.linear_fc2.weight
module.decoder.layers.86.mixer.norm.weight
module.decoder.layers.85.mlp.linear_fc2.weight
module.decoder.layers.84.mixer.conv1d.bias
module.decoder.layers.84.mixer.conv1d.weight
module.decoder.layers.86.mixer.in_proj.layer_norm_weight
module.decoder.layers.86.mixer.conv1d.weight
module.decoder.layers.85.mlp.linear_fc1.weight
module.decoder.layers.88.mixer.in_proj.layer_norm_weight
module.decoder.layers.87.mlp.linear_fc1.weight
module.decoder.layers.86.mixer.out_proj.weight
module.decoder.layers.85.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.84.mixer.out_proj.weight
module.decoder.layers.84.mixer.in_proj.weight
module.decoder.layers.86.mixer.dt_bias
module.decoder.layers.88.mixer.dt_bias
module.decoder.layers.86.mixer.conv1d.bias
module.decoder.layers.86.mixer.A_log
module.decoder.layers.88.mixer.A_log
module.decoder.layers.88.mixer.D
module.decoder.layers.87.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.86.mixer.D
module.decoder.layers.84.mixer.norm.weight
Params for bucket 9 (41342874 elements, 41343232 padded size):
module.decoder.layers.82.mixer.dt_bias
module.decoder.layers.81.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.81.mlp.linear_fc2.weight
module.decoder.layers.80.self_attention.linear_qkv.layer_norm_weight
module.decoder.layers.82.mixer.D
module.decoder.layers.84.mixer.in_proj.layer_norm_weight
module.decoder.layers.84.mixer.D
module.decoder.layers.83.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.82.mixer.norm.weight
module.decoder.layers.82.mixer.A_log
module.decoder.layers.80.self_attention.linear_qkv.weight
module.decoder.layers.82.mixer.in_proj.weight
module.decoder.layers.83.mlp.linear_fc2.weight
module.decoder.layers.84.mixer.dt_bias
module.decoder.layers.82.mixer.conv1d.bias
module.decoder.layers.80.self_attention.linear_qkv.bias
module.decoder.layers.80.self_attention.linear_proj.weight
module.decoder.layers.82.mixer.conv1d.weight
module.decoder.layers.82.mixer.in_proj.layer_norm_weight
module.decoder.layers.81.mlp.linear_fc1.weight
module.decoder.layers.84.mixer.A_log
module.decoder.layers.83.mlp.linear_fc1.weight
module.decoder.layers.82.mixer.out_proj.weight
Params for bucket 10 (46174125 elements, 46174336 padded size):
module.decoder.layers.78.mixer.A_log
module.decoder.layers.77.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.79.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.78.mixer.in_proj.layer_norm_weight
module.decoder.layers.78.mixer.D
module.decoder.layers.77.mlp.linear_fc2.weight
module.decoder.layers.79.mlp.linear_fc2.weight
module.decoder.layers.78.mixer.in_proj.weight
module.decoder.layers.76.mixer.in_proj.weight
module.decoder.layers.76.mixer.conv1d.weight
module.decoder.layers.76.mixer.out_proj.weight
module.decoder.layers.79.mlp.linear_fc1.weight
module.decoder.layers.78.mixer.norm.weight
module.decoder.layers.78.mixer.out_proj.weight
module.decoder.layers.78.mixer.conv1d.weight
module.decoder.layers.77.mlp.linear_fc1.weight
module.decoder.layers.76.mixer.norm.weight
module.decoder.layers.76.mixer.conv1d.bias
module.decoder.layers.78.mixer.conv1d.bias
module.decoder.layers.78.mixer.dt_bias
Params for bucket 11 (46176090 elements, 46176384 padded size):
module.decoder.layers.74.mixer.D
module.decoder.layers.74.mixer.A_log
module.decoder.layers.73.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.76.mixer.in_proj.layer_norm_weight
module.decoder.layers.76.mixer.D
module.decoder.layers.76.mixer.A_log
module.decoder.layers.75.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.74.mixer.in_proj.layer_norm_weight
module.decoder.layers.73.mlp.linear_fc2.weight
module.decoder.layers.75.mlp.linear_fc2.weight
module.decoder.layers.74.mixer.in_proj.weight
module.decoder.layers.72.mixer.in_proj.weight
module.decoder.layers.73.mlp.linear_fc1.weight
module.decoder.layers.75.mlp.linear_fc1.weight
module.decoder.layers.74.mixer.out_proj.weight
module.decoder.layers.74.mixer.norm.weight
module.decoder.layers.74.mixer.conv1d.weight
module.decoder.layers.72.mixer.out_proj.weight
module.decoder.layers.72.mixer.norm.weight
module.decoder.layers.72.mixer.conv1d.weight
module.decoder.layers.76.mixer.dt_bias
module.decoder.layers.74.mixer.conv1d.bias
module.decoder.layers.74.mixer.dt_bias
module.decoder.layers.72.mixer.conv1d.bias
Params for bucket 12 (46176090 elements, 46176384 padded size):
module.decoder.layers.70.mixer.D
module.decoder.layers.70.mixer.A_log
module.decoder.layers.69.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.72.mixer.in_proj.layer_norm_weight
module.decoder.layers.72.mixer.D
module.decoder.layers.72.mixer.A_log
module.decoder.layers.71.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.70.mixer.in_proj.layer_norm_weight
module.decoder.layers.69.mlp.linear_fc2.weight
module.decoder.layers.71.mlp.linear_fc2.weight
module.decoder.layers.70.mixer.in_proj.weight
module.decoder.layers.68.mixer.in_proj.weight
module.decoder.layers.69.mlp.linear_fc1.weight
module.decoder.layers.71.mlp.linear_fc1.weight
module.decoder.layers.70.mixer.out_proj.weight
module.decoder.layers.70.mixer.norm.weight
module.decoder.layers.70.mixer.conv1d.weight
module.decoder.layers.68.mixer.out_proj.weight
module.decoder.layers.68.mixer.norm.weight
module.decoder.layers.68.mixer.conv1d.weight
module.decoder.layers.72.mixer.dt_bias
module.decoder.layers.70.mixer.conv1d.bias
module.decoder.layers.70.mixer.dt_bias
module.decoder.layers.68.mixer.conv1d.bias
Params for bucket 13 (41342874 elements, 41343232 padded size):
module.decoder.layers.66.mixer.A_log
module.decoder.layers.68.mixer.in_proj.layer_norm_weight
module.decoder.layers.68.mixer.D
module.decoder.layers.68.mixer.A_log
module.decoder.layers.67.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.66.mixer.norm.weight
module.decoder.layers.66.mixer.D
module.decoder.layers.64.self_attention.linear_qkv.weight
module.decoder.layers.65.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.67.mlp.linear_fc2.weight
module.decoder.layers.66.mixer.conv1d.bias
module.decoder.layers.66.mixer.in_proj.weight
module.decoder.layers.64.self_attention.linear_qkv.bias
module.decoder.layers.64.self_attention.linear_proj.weight
module.decoder.layers.65.mlp.linear_fc1.weight
module.decoder.layers.67.mlp.linear_fc1.weight
module.decoder.layers.66.mixer.out_proj.weight
module.decoder.layers.66.mixer.conv1d.weight
module.decoder.layers.66.mixer.in_proj.layer_norm_weight
module.decoder.layers.65.mlp.linear_fc2.weight
module.decoder.layers.66.mixer.dt_bias
module.decoder.layers.68.mixer.dt_bias
module.decoder.layers.64.self_attention.linear_qkv.layer_norm_weight
Params for bucket 14 (46174125 elements, 46174336 padded size):
module.decoder.layers.61.mlp.linear_fc1.weight
module.decoder.layers.63.mlp.linear_fc2.weight
module.decoder.layers.62.mixer.in_proj.weight
module.decoder.layers.62.mixer.in_proj.layer_norm_weight
module.decoder.layers.60.mixer.conv1d.bias
module.decoder.layers.61.mlp.linear_fc2.weight
module.decoder.layers.63.mlp.linear_fc1.weight
module.decoder.layers.62.mixer.out_proj.weight
module.decoder.layers.62.mixer.conv1d.bias
module.decoder.layers.62.mixer.conv1d.weight
module.decoder.layers.62.mixer.D
module.decoder.layers.61.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.62.mixer.dt_bias
module.decoder.layers.60.mixer.in_proj.weight
module.decoder.layers.60.mixer.out_proj.weight
module.decoder.layers.60.mixer.norm.weight
module.decoder.layers.62.mixer.norm.weight
module.decoder.layers.63.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.62.mixer.A_log
module.decoder.layers.60.mixer.conv1d.weight
Params for bucket 15 (46176090 elements, 46176384 padded size):
module.decoder.layers.58.mixer.dt_bias
module.decoder.layers.60.mixer.dt_bias
module.decoder.layers.58.mixer.conv1d.bias
module.decoder.layers.56.mixer.conv1d.bias
module.decoder.layers.58.mixer.in_proj.layer_norm_weight
module.decoder.layers.58.mixer.D
module.decoder.layers.60.mixer.in_proj.layer_norm_weight
module.decoder.layers.60.mixer.A_log
module.decoder.layers.60.mixer.D
module.decoder.layers.59.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.58.mixer.A_log
module.decoder.layers.57.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.57.mlp.linear_fc2.weight
module.decoder.layers.58.mixer.in_proj.weight
module.decoder.layers.59.mlp.linear_fc2.weight
module.decoder.layers.56.mixer.in_proj.weight
module.decoder.layers.58.mixer.conv1d.weight
module.decoder.layers.57.mlp.linear_fc1.weight
module.decoder.layers.59.mlp.linear_fc1.weight
module.decoder.layers.58.mixer.out_proj.weight
module.decoder.layers.58.mixer.norm.weight
module.decoder.layers.56.mixer.out_proj.weight
module.decoder.layers.56.mixer.norm.weight
module.decoder.layers.56.mixer.conv1d.weight
Params for bucket 16 (46176090 elements, 46176384 padded size):
module.decoder.layers.54.mixer.dt_bias
module.decoder.layers.56.mixer.dt_bias
module.decoder.layers.54.mixer.conv1d.bias
module.decoder.layers.52.mixer.conv1d.bias
module.decoder.layers.54.mixer.in_proj.layer_norm_weight
module.decoder.layers.54.mixer.D
module.decoder.layers.56.mixer.in_proj.layer_norm_weight
module.decoder.layers.56.mixer.A_log
module.decoder.layers.56.mixer.D
module.decoder.layers.55.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.54.mixer.A_log
module.decoder.layers.53.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.54.mixer.in_proj.weight
module.decoder.layers.55.mlp.linear_fc2.weight
module.decoder.layers.53.mlp.linear_fc2.weight
module.decoder.layers.52.mixer.in_proj.weight
module.decoder.layers.54.mixer.conv1d.weight
module.decoder.layers.53.mlp.linear_fc1.weight
module.decoder.layers.55.mlp.linear_fc1.weight
module.decoder.layers.54.mixer.out_proj.weight
module.decoder.layers.54.mixer.norm.weight
module.decoder.layers.52.mixer.out_proj.weight
module.decoder.layers.52.mixer.norm.weight
module.decoder.layers.52.mixer.conv1d.weight
Params for bucket 17 (41342874 elements, 41343232 padded size):
module.decoder.layers.50.mixer.dt_bias
module.decoder.layers.52.mixer.dt_bias
module.decoder.layers.49.mlp.linear_fc2.weight
module.decoder.layers.48.self_attention.linear_qkv.layer_norm_weight
module.decoder.layers.50.mixer.D
module.decoder.layers.52.mixer.in_proj.layer_norm_weight
module.decoder.layers.52.mixer.A_log
module.decoder.layers.52.mixer.D
module.decoder.layers.51.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.50.mixer.norm.weight
module.decoder.layers.50.mixer.A_log
module.decoder.layers.48.self_attention.linear_qkv.weight
module.decoder.layers.50.mixer.in_proj.weight
module.decoder.layers.49.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.51.mlp.linear_fc2.weight
module.decoder.layers.50.mixer.conv1d.bias
module.decoder.layers.48.self_attention.linear_qkv.bias
module.decoder.layers.48.self_attention.linear_proj.weight
module.decoder.layers.50.mixer.conv1d.weight
module.decoder.layers.50.mixer.in_proj.layer_norm_weight
module.decoder.layers.51.mlp.linear_fc1.weight
module.decoder.layers.50.mixer.out_proj.weight
module.decoder.layers.49.mlp.linear_fc1.weight
Params for bucket 18 (46174125 elements, 46174336 padded size):
module.decoder.layers.46.mixer.A_log
module.decoder.layers.45.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.47.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.46.mixer.D
module.decoder.layers.45.mlp.linear_fc2.weight
module.decoder.layers.44.mixer.norm.weight
module.decoder.layers.47.mlp.linear_fc2.weight
module.decoder.layers.46.mixer.norm.weight
module.decoder.layers.46.mixer.in_proj.weight
module.decoder.layers.44.mixer.in_proj.weight
module.decoder.layers.44.mixer.out_proj.weight
module.decoder.layers.45.mlp.linear_fc1.weight
module.decoder.layers.46.mixer.out_proj.weight
module.decoder.layers.47.mlp.linear_fc1.weight
module.decoder.layers.46.mixer.conv1d.weight
module.decoder.layers.46.mixer.in_proj.layer_norm_weight
module.decoder.layers.44.mixer.conv1d.weight
module.decoder.layers.44.mixer.conv1d.bias
module.decoder.layers.46.mixer.conv1d.bias
module.decoder.layers.46.mixer.dt_bias
Params for bucket 19 (46176090 elements, 46176384 padded size):
module.decoder.layers.42.mixer.A_log
module.decoder.layers.44.mixer.A_log
module.decoder.layers.44.mixer.D
module.decoder.layers.43.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.42.mixer.D
module.decoder.layers.41.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.40.mixer.norm.weight
module.decoder.layers.41.mlp.linear_fc2.weight
module.decoder.layers.43.mlp.linear_fc2.weight
module.decoder.layers.42.mixer.norm.weight
module.decoder.layers.42.mixer.in_proj.weight
module.decoder.layers.40.mixer.in_proj.weight
module.decoder.layers.42.mixer.in_proj.layer_norm_weight
module.decoder.layers.42.mixer.conv1d.weight
module.decoder.layers.44.mixer.in_proj.layer_norm_weight
module.decoder.layers.43.mlp.linear_fc1.weight
module.decoder.layers.42.mixer.out_proj.weight
module.decoder.layers.41.mlp.linear_fc1.weight
module.decoder.layers.40.mixer.out_proj.weight
module.decoder.layers.40.mixer.conv1d.bias
module.decoder.layers.40.mixer.conv1d.weight
module.decoder.layers.44.mixer.dt_bias
module.decoder.layers.42.mixer.conv1d.bias
module.decoder.layers.42.mixer.dt_bias
Params for bucket 20 (46176090 elements, 46176384 padded size):
module.decoder.layers.38.mixer.conv1d.weight
module.decoder.layers.40.mixer.A_log
module.decoder.layers.38.mixer.out_proj.weight
module.decoder.layers.38.mixer.norm.weight
module.decoder.layers.37.mlp.linear_fc1.weight
module.decoder.layers.36.mixer.out_proj.weight
module.decoder.layers.36.mixer.norm.weight
module.decoder.layers.36.mixer.conv1d.weight
module.decoder.layers.40.mixer.in_proj.layer_norm_weight
module.decoder.layers.39.mlp.linear_fc1.weight
module.decoder.layers.38.mixer.conv1d.bias
module.decoder.layers.38.mixer.dt_bias
module.decoder.layers.36.mixer.conv1d.bias
module.decoder.layers.38.mixer.in_proj.layer_norm_weight
module.decoder.layers.37.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.38.mixer.A_log
module.decoder.layers.40.mixer.D
module.decoder.layers.39.mlp.linear_fc2.weight
module.decoder.layers.39.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.38.mixer.D
module.decoder.layers.37.mlp.linear_fc2.weight
module.decoder.layers.40.mixer.dt_bias
module.decoder.layers.38.mixer.in_proj.weight
module.decoder.layers.36.mixer.in_proj.weight
Params for bucket 21 (41342874 elements, 41343232 padded size):
module.decoder.layers.33.mlp.linear_fc1.weight
module.decoder.layers.34.mixer.conv1d.weight
module.decoder.layers.35.mlp.linear_fc1.weight
module.decoder.layers.34.mixer.out_proj.weight
module.decoder.layers.34.mixer.in_proj.layer_norm_weight
module.decoder.layers.36.mixer.dt_bias
module.decoder.layers.34.mixer.dt_bias
module.decoder.layers.33.mlp.linear_fc2.weight
module.decoder.layers.32.self_attention.linear_qkv.layer_norm_weight
module.decoder.layers.34.mixer.D
module.decoder.layers.36.mixer.in_proj.layer_norm_weight
module.decoder.layers.36.mixer.D
module.decoder.layers.36.mixer.A_log
module.decoder.layers.35.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.34.mixer.norm.weight
module.decoder.layers.34.mixer.A_log
module.decoder.layers.32.self_attention.linear_qkv.weight
module.decoder.layers.33.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.35.mlp.linear_fc2.weight
module.decoder.layers.34.mixer.conv1d.bias
module.decoder.layers.34.mixer.in_proj.weight
module.decoder.layers.32.self_attention.linear_qkv.bias
module.decoder.layers.32.self_attention.linear_proj.weight
Params for bucket 22 (46174125 elements, 46174336 padded size):
module.decoder.layers.30.mixer.dt_bias
module.decoder.layers.30.mixer.conv1d.bias
module.decoder.layers.28.mixer.conv1d.bias
module.decoder.layers.29.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.30.mixer.A_log
module.decoder.layers.31.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.30.mixer.D
module.decoder.layers.28.mixer.norm.weight
module.decoder.layers.31.mlp.linear_fc2.weight
module.decoder.layers.30.mixer.norm.weight
module.decoder.layers.30.mixer.in_proj.weight
module.decoder.layers.29.mlp.linear_fc2.weight
module.decoder.layers.28.mixer.in_proj.weight
module.decoder.layers.29.mlp.linear_fc1.weight
module.decoder.layers.28.mixer.out_proj.weight
module.decoder.layers.28.mixer.conv1d.weight
module.decoder.layers.30.mixer.out_proj.weight
module.decoder.layers.31.mlp.linear_fc1.weight
module.decoder.layers.30.mixer.conv1d.weight
module.decoder.layers.30.mixer.in_proj.layer_norm_weight
Params for bucket 23 (46176090 elements, 46176384 padded size):
module.decoder.layers.26.mixer.dt_bias
module.decoder.layers.28.mixer.dt_bias
module.decoder.layers.26.mixer.conv1d.bias
module.decoder.layers.24.mixer.conv1d.bias
module.decoder.layers.25.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.26.mixer.D
module.decoder.layers.26.mixer.A_log
module.decoder.layers.28.mixer.D
module.decoder.layers.28.mixer.A_log
module.decoder.layers.27.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.27.mlp.linear_fc2.weight
module.decoder.layers.26.mixer.norm.weight
module.decoder.layers.25.mlp.linear_fc2.weight
module.decoder.layers.26.mixer.in_proj.weight
module.decoder.layers.24.mixer.norm.weight
module.decoder.layers.24.mixer.in_proj.weight
module.decoder.layers.25.mlp.linear_fc1.weight
module.decoder.layers.28.mixer.in_proj.layer_norm_weight
module.decoder.layers.27.mlp.linear_fc1.weight
module.decoder.layers.26.mixer.out_proj.weight
module.decoder.layers.26.mixer.conv1d.weight
module.decoder.layers.26.mixer.in_proj.layer_norm_weight
module.decoder.layers.24.mixer.out_proj.weight
module.decoder.layers.24.mixer.conv1d.weight
Params for bucket 24 (46176090 elements, 46176384 padded size):
module.decoder.layers.22.mixer.dt_bias
module.decoder.layers.24.mixer.dt_bias
module.decoder.layers.22.mixer.conv1d.bias
module.decoder.layers.20.mixer.conv1d.bias
module.decoder.layers.21.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.22.mixer.D
module.decoder.layers.22.mixer.A_log
module.decoder.layers.24.mixer.D
module.decoder.layers.24.mixer.A_log
module.decoder.layers.23.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.23.mlp.linear_fc2.weight
module.decoder.layers.22.mixer.norm.weight
module.decoder.layers.21.mlp.linear_fc2.weight
module.decoder.layers.22.mixer.in_proj.weight
module.decoder.layers.20.mixer.norm.weight
module.decoder.layers.20.mixer.in_proj.weight
module.decoder.layers.21.mlp.linear_fc1.weight
module.decoder.layers.24.mixer.in_proj.layer_norm_weight
module.decoder.layers.23.mlp.linear_fc1.weight
module.decoder.layers.22.mixer.out_proj.weight
module.decoder.layers.22.mixer.in_proj.layer_norm_weight
module.decoder.layers.22.mixer.conv1d.weight
module.decoder.layers.20.mixer.out_proj.weight
module.decoder.layers.20.mixer.conv1d.weight
Params for bucket 25 (41342874 elements, 41343232 padded size):
module.decoder.layers.18.mixer.dt_bias
module.decoder.layers.20.mixer.dt_bias
module.decoder.layers.17.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.16.self_attention.linear_qkv.bias
module.decoder.layers.16.self_attention.linear_proj.weight
module.decoder.layers.18.mixer.A_log
module.decoder.layers.20.mixer.D
module.decoder.layers.20.mixer.A_log
module.decoder.layers.19.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.17.mlp.linear_fc1.weight
module.decoder.layers.18.mixer.in_proj.layer_norm_weight
module.decoder.layers.17.mlp.linear_fc2.weight
module.decoder.layers.19.mlp.linear_fc2.weight
module.decoder.layers.18.mixer.conv1d.bias
module.decoder.layers.18.mixer.in_proj.weight
module.decoder.layers.16.self_attention.linear_qkv.layer_norm_weight
module.decoder.layers.18.mixer.D
module.decoder.layers.20.mixer.in_proj.layer_norm_weight
module.decoder.layers.19.mlp.linear_fc1.weight
module.decoder.layers.18.mixer.out_proj.weight
module.decoder.layers.18.mixer.norm.weight
module.decoder.layers.18.mixer.conv1d.weight
module.decoder.layers.16.self_attention.linear_qkv.weight
Params for bucket 26 (46174125 elements, 46174336 padded size):
module.decoder.layers.13.mlp.linear_fc1.weight
module.decoder.layers.12.mixer.conv1d.weight
module.decoder.layers.15.mlp.linear_fc1.weight
module.decoder.layers.14.mixer.out_proj.weight
module.decoder.layers.14.mixer.conv1d.weight
module.decoder.layers.14.mixer.in_proj.layer_norm_weight
module.decoder.layers.12.mixer.out_proj.weight
module.decoder.layers.14.mixer.conv1d.bias
module.decoder.layers.14.mixer.dt_bias
module.decoder.layers.12.mixer.conv1d.bias
module.decoder.layers.15.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.14.mixer.D
module.decoder.layers.14.mixer.A_log
module.decoder.layers.13.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.13.mlp.linear_fc2.weight
module.decoder.layers.12.mixer.norm.weight
module.decoder.layers.15.mlp.linear_fc2.weight
module.decoder.layers.14.mixer.norm.weight
module.decoder.layers.14.mixer.in_proj.weight
module.decoder.layers.12.mixer.in_proj.weight
Params for bucket 27 (46176090 elements, 46176384 padded size):
module.decoder.layers.10.mixer.conv1d.weight
module.decoder.layers.9.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.12.mixer.in_proj.layer_norm_weight
module.decoder.layers.11.mlp.linear_fc1.weight
module.decoder.layers.10.mixer.out_proj.weight
module.decoder.layers.10.mixer.norm.weight
module.decoder.layers.9.mlp.linear_fc2.weight
module.decoder.layers.12.mixer.dt_bias
module.decoder.layers.10.mixer.dt_bias
module.decoder.layers.8.mixer.norm.weight
module.decoder.layers.8.mixer.in_proj.weight
module.decoder.layers.9.mlp.linear_fc1.weight
module.decoder.layers.12.mixer.A_log
module.decoder.layers.12.mixer.D
module.decoder.layers.10.mixer.conv1d.bias
module.decoder.layers.10.mixer.A_log
module.decoder.layers.10.mixer.D
module.decoder.layers.8.mixer.out_proj.weight
module.decoder.layers.8.mixer.conv1d.weight
module.decoder.layers.10.mixer.in_proj.layer_norm_weight
module.decoder.layers.10.mixer.in_proj.weight
module.decoder.layers.11.mlp.linear_fc2.weight
module.decoder.layers.11.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.8.mixer.conv1d.bias
Params for bucket 28 (46176090 elements, 46176384 padded size):
module.decoder.layers.5.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.8.mixer.D
module.decoder.layers.8.mixer.A_log
module.decoder.layers.7.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.6.mixer.D
module.decoder.layers.6.mixer.A_log
module.decoder.layers.6.mixer.in_proj.weight
module.decoder.layers.5.mlp.linear_fc2.weight
module.decoder.layers.7.mlp.linear_fc2.weight
module.decoder.layers.6.mixer.norm.weight
module.decoder.layers.4.mixer.conv1d.bias
module.decoder.layers.4.mixer.in_proj.weight
module.decoder.layers.6.mixer.conv1d.weight
module.decoder.layers.5.mlp.linear_fc1.weight
module.decoder.layers.8.mixer.in_proj.layer_norm_weight
module.decoder.layers.7.mlp.linear_fc1.weight
module.decoder.layers.6.mixer.out_proj.weight
module.decoder.layers.6.mixer.in_proj.layer_norm_weight
module.decoder.layers.4.mixer.out_proj.weight
module.decoder.layers.4.mixer.norm.weight
module.decoder.layers.4.mixer.conv1d.weight
module.decoder.layers.6.mixer.dt_bias
module.decoder.layers.8.mixer.dt_bias
module.decoder.layers.6.mixer.conv1d.bias
Params for bucket 29 (41342874 elements, 41343232 padded size):
module.decoder.layers.2.mixer.dt_bias
module.decoder.layers.4.mixer.A_log
module.decoder.layers.3.mlp.linear_fc1.weight
module.decoder.layers.0.self_attention.linear_qkv.weight
module.decoder.layers.0.self_attention.linear_proj.weight
module.decoder.layers.4.mixer.in_proj.layer_norm_weight
module.decoder.layers.3.mlp.linear_fc1.layer_norm_weight
module.decoder.layers.2.mixer.A_log
module.decoder.layers.0.self_attention.linear_qkv.bias
module.decoder.layers.2.mixer.in_proj.weight
module.decoder.layers.1.mlp.linear_fc2.weight
module.decoder.layers.2.mixer.D
module.decoder.layers.3.mlp.linear_fc2.weight
module.decoder.layers.2.mixer.norm.weight
module.decoder.layers.2.mixer.conv1d.bias
module.decoder.layers.0.self_attention.linear_qkv.layer_norm_weight
module.decoder.layers.2.mixer.conv1d.weight
module.decoder.layers.1.mlp.linear_fc1.weight
module.decoder.layers.4.mixer.dt_bias
module.decoder.layers.4.mixer.D
module.decoder.layers.2.mixer.out_proj.weight
module.decoder.layers.2.mixer.in_proj.layer_norm_weight
module.decoder.layers.1.mlp.linear_fc1.layer_norm_weight
Params for bucket 30 (95109120 elements, 95109120 padded size):
module.embedding.word_embeddings.weight
[MambaModel(
(embedding): LanguageModelEmbedding(
(word_embeddings): VocabParallelEmbedding()
(embedding_dropout): Dropout(p=0.0, inplace=False)
)
(rotary_pos_emb): RotaryEmbedding()
(decoder): MambaStack(
(layers): ModuleList(
(0): TransformerLayer(
(input_layernorm): IdentityOp()
(self_attention): SelfAttention(
(core_attention): TEDotProductAttention(
(flash_attention): FlashAttention()
(fused_attention): FusedAttention()
(unfused_attention): UnfusedDotProductAttention(
(scale_mask_softmax): FusedScaleMaskSoftmax()
(attention_dropout): Dropout(p=0.0, inplace=False)
)
)
(linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
(linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2)
(q_layernorm): IdentityOp()
(k_layernorm): IdentityOp()
)
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): IdentityOp()
(mlp_bda): IdentityFuncOp()
)
(1): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(2): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(3): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(4): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(5): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(6): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(7): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(8): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(9): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(10): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(11): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(12): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(13): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(14): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(15): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(16): TransformerLayer(
(input_layernorm): IdentityOp()
(self_attention): SelfAttention(
(core_attention): TEDotProductAttention(
(flash_attention): FlashAttention()
(fused_attention): FusedAttention()
(unfused_attention): UnfusedDotProductAttention(
(scale_mask_softmax): FusedScaleMaskSoftmax()
(attention_dropout): Dropout(p=0.0, inplace=False)
)
)
(linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
(linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2)
(q_layernorm): IdentityOp()
(k_layernorm): IdentityOp()
)
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): IdentityOp()
(mlp_bda): IdentityFuncOp()
)
(17): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(18): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(19): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(20): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(21): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(22): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(23): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(24): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(25): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(26): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(27): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(28): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(29): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(30): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(31): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(32): TransformerLayer(
(input_layernorm): IdentityOp()
(self_attention): SelfAttention(
(core_attention): TEDotProductAttention(
(flash_attention): FlashAttention()
(fused_attention): FusedAttention()
(unfused_attention): UnfusedDotProductAttention(
(scale_mask_softmax): FusedScaleMaskSoftmax()
(attention_dropout): Dropout(p=0.0, inplace=False)
)
)
(linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
(linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2)
(q_layernorm): IdentityOp()
(k_layernorm): IdentityOp()
)
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): IdentityOp()
(mlp_bda): IdentityFuncOp()
)
(33): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(34): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(35): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(36): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(37): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(38): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(39): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(40): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(41): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(42): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(43): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(44): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(45): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(46): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(47): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(48): TransformerLayer(
(input_layernorm): IdentityOp()
(self_attention): SelfAttention(
(core_attention): TEDotProductAttention(
(flash_attention): FlashAttention()
(fused_attention): FusedAttention()
(unfused_attention): UnfusedDotProductAttention(
(scale_mask_softmax): FusedScaleMaskSoftmax()
(attention_dropout): Dropout(p=0.0, inplace=False)
)
)
(linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
(linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2)
(q_layernorm): IdentityOp()
(k_layernorm): IdentityOp()
)
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): IdentityOp()
(mlp_bda): IdentityFuncOp()
)
(49): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(50): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(51): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(52): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(53): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(54): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(55): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(56): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(57): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(58): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(59): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(60): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(61): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(62): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(63): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(64): TransformerLayer(
(input_layernorm): IdentityOp()
(self_attention): SelfAttention(
(core_attention): TEDotProductAttention(
(flash_attention): FlashAttention()
(fused_attention): FusedAttention()
(unfused_attention): UnfusedDotProductAttention(
(scale_mask_softmax): FusedScaleMaskSoftmax()
(attention_dropout): Dropout(p=0.0, inplace=False)
)
)
(linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
(linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2)
(q_layernorm): IdentityOp()
(k_layernorm): IdentityOp()
)
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): IdentityOp()
(mlp_bda): IdentityFuncOp()
)
(65): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(66): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(67): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(68): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(69): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(70): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(71): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(72): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(73): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(74): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(75): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(76): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(77): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(78): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(79): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(80): TransformerLayer(
(input_layernorm): IdentityOp()
(self_attention): SelfAttention(
(core_attention): TEDotProductAttention(
(flash_attention): FlashAttention()
(fused_attention): FusedAttention()
(unfused_attention): UnfusedDotProductAttention(
(scale_mask_softmax): FusedScaleMaskSoftmax()
(attention_dropout): Dropout(p=0.0, inplace=False)
)
)
(linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
(linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2)
(q_layernorm): IdentityOp()
(k_layernorm): IdentityOp()
)
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): IdentityOp()
(mlp_bda): IdentityFuncOp()
)
(81): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(82): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(83): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(84): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(85): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(86): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(87): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(88): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(89): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(90): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(91): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(92): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(93): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(94): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(95): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(96): TransformerLayer(
(input_layernorm): IdentityOp()
(self_attention): SelfAttention(
(core_attention): TEDotProductAttention(
(flash_attention): FlashAttention()
(fused_attention): FusedAttention()
(unfused_attention): UnfusedDotProductAttention(
(scale_mask_softmax): FusedScaleMaskSoftmax()
(attention_dropout): Dropout(p=0.0, inplace=False)
)
)
(linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
(linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2)
(q_layernorm): IdentityOp()
(k_layernorm): IdentityOp()
)
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): IdentityOp()
(mlp_bda): IdentityFuncOp()
)
(97): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(98): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(99): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(100): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(101): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(102): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(103): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(104): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(105): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(106): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(107): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(108): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(109): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
(110): MambaLayer(
(mixer): MambaMixer(
(in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2)
(conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880)
(act): SiLU()
(norm): ExtendedRMSNorm()
(out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2)
)
(norm): IdentityOp()
)
(111): MLPLayer(
(input_layernorm): IdentityOp()
(self_attention): IdentityOp()
(self_attn_bda): IdentityFuncOp()
(pre_cross_attn_layernorm): IdentityOp()
(cross_attention): IdentityOp()
(cross_attn_bda): IdentityFuncOp()
(pre_mlp_layernorm): IdentityOp()
(mlp): MLP(
(linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2)
(linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2)
)
)
)
(final_norm): RMSNorm()
)
(output_layer): ColumnParallelLinear(in_features=1920, out_features=99072, bias=False, TP=2)
)]
[DEBUG] freeze_non_mamba: False
INFO:megatron.core.optimizer:Setting up optimizer with config OptimizerConfig(optimizer='adam', lr=2e-05, min_lr=7e-07, decoupled_lr=None, decoupled_min_lr=None, weight_decay=0.1, fp16=False, bf16=True, params_dtype=torch.bfloat16, use_precision_aware_optimizer=False, main_grads_dtype=torch.float32, main_params_dtype=torch.float32, exp_avg_dtype=torch.float32, exp_avg_sq_dtype=torch.float32, loss_scale=None, initial_loss_scale=4294967296, min_loss_scale=1.0, loss_scale_window=1000, hysteresis=2, adam_beta1=0.9, adam_beta2=0.999, adam_eps=1e-08, sgd_momentum=0.9, muon_momentum=0.95, muon_nesterov=True, muon_ns_steps=5, muon_matched_adamw_rms=0.2, use_distributed_optimizer=True, overlap_param_gather_with_optimizer_step=False, optimizer_cpu_offload=False, optimizer_offload_fraction=1.0, use_torch_optimizer_for_cpu_offload=False, overlap_cpu_optimizer_d2h_h2d=False, pin_cpu_grads=True, pin_cpu_params=True, clip_grad=0.5, log_num_zeros_in_grad=False, barrier_with_L1_time=True, timers=<megatron.core.timers.Timers object at 0x7efd9ec4c350>, config_logger_dir='')
setting training iterations to 59
INFO:megatron.core.optimizer_param_scheduler:> learning rate decay style: linear
[DEBUG] freeze_non_mamba: False
[DEBUG] freeze_non_mamba: False
[DEBUG] freeze_non_mamba: False
[DEBUG] freeze_non_mamba: False
loading checkpoint from /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/RADLADS-paper/out/L56-D1920-qwen_mamba2_qwen2-e1-i1920-s320-hd64-gn6-A0-S512--step1-dclm10b/rwkv-394-hf-A7-0_8_16_24_32_40_48/megatron-pp1-tp2 at iteration 0
could not find arguments in the checkpoint ...
checkpoint version 0
successfully fixed query-key-values ordering for checkpoint version 0
successfully loaded checkpoint from /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/RADLADS-paper/out/L56-D1920-qwen_mamba2_qwen2-e1-i1920-s320-hd64-gn6-A0-S512--step1-dclm10b/rwkv-394-hf-A7-0_8_16_24_32_40_48/megatron-pp1-tp2 [ t 1/2, p 1/1 ] at iteration 0
(min, max) time across ranks (ms):
load-checkpoint ................................: (6511.67, 6511.73)
[after model, optimizer, and learning rate scheduler are built] datetime: 2025-09-17 22:26:16
saving checkpoint at iteration 0 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 in torch format
successfully saved checkpoint from iteration 0 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 [ t 1/2, p 1/1 ]
> building train, validation, and test datasets ...
> datasets target sizes (minimum size):
train: 61035
validation: 10240
test: 10240
INFO:megatron.core.datasets.blended_megatron_dataset_config:Let split_matrix = [(0, 1.0), None, None]
> building train, validation, and test datasets for GPT ...
INFO:megatron.core.datasets.blended_megatron_dataset_builder:Building GPTDataset splits with sizes=(61035, 10240, 10240) and config=GPTDatasetConfig(random_seed=1234, sequence_length=32768, blend=(['/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache/datasets/huggingface/Teaven/combine_2B_0908/binidx/yulan_mini'], None), blend_per_split=None, split='100,0,0', split_matrix=[(0, 1.0), None, None], num_dataset_builder_threads=1, path_to_cache='/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache', mmap_bin_files=True, mock=False, tokenizer=<megatron.training.tokenizer.tokenizer._HuggingFaceTokenizer object at 0x7efe4d2af6d0>, mid_level_dataset_surplus=0.005, reset_position_ids=False, reset_attention_mask=False, eod_mask_loss=False, create_attention_mask=False, drop_last_partial_validation_sequence=True, add_extra_token_to_sequence=True, object_storage_cache_path=None)
INFO:megatron.core.datasets.indexed_dataset:Load the _IndexReader from /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache/datasets/huggingface/Teaven/combine_2B_0908/binidx/yulan_mini.idx
INFO:megatron.core.datasets.indexed_dataset: Extract the sequence lengths
INFO:megatron.core.datasets.indexed_dataset: Extract the sequence pointers
INFO:megatron.core.datasets.indexed_dataset: Extract the document indices
INFO:megatron.core.datasets.indexed_dataset:> total number of sequences: 73736
INFO:megatron.core.datasets.indexed_dataset:> total number of documents: 73736
INFO:megatron.core.datasets.gpt_dataset:Load the GPTDataset train indices
INFO:megatron.core.datasets.gpt_dataset: Load the document index from 195a986bb6005248efb3de26e9334e88-GPTDataset-train-document_index.npy
INFO:megatron.core.datasets.gpt_dataset: Load the sample index from 195a986bb6005248efb3de26e9334e88-GPTDataset-train-sample_index.npy
INFO:megatron.core.datasets.gpt_dataset: Load the shuffle index from 195a986bb6005248efb3de26e9334e88-GPTDataset-train-shuffle_index.npy
INFO:megatron.core.datasets.gpt_dataset:> total number of samples: 64518
> finished creating GPT datasets ...
[after dataloaders are built] datetime: 2025-09-17 22:26:22
done with setup ...
(min, max) time across ranks (ms):
model-and-optimizer-setup ......................: (8145.52, 8201.88)
train/valid/test-data-iterators-setup ..........: (12.45, 100.39)
training ...
Setting rerun_state_machine.current_iteration to 0...
[before the start of training step] datetime: 2025-09-17 22:26:22
Iter 0 increase timing-log-level to 2.
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
[Rank 2] (after 1 iterations) memory (MB) | allocated: 12648.54443359375 | max allocated: 44991.1953125 | reserved: 47306.0 | max reserved: 47306.0 | MEM: 48.63%
[Rank 3] (after 1 iterations) memory (MB) | allocated: 12648.54443359375 | max allocated: 44991.1953125 | reserved: 47306.0 | max reserved: 47306.0 | MEM: 48.63%
[Rank 1] (after 1 iterations) memory (MB) | allocated: 12648.5341796875 | max allocated: 44991.19189453125 | reserved: 47224.0 | max reserved: 47224.0 | MEM: 48.54%
[2025-09-17 22:44:07] iteration 1/ 59 | consumed samples: 1024 | elapsed time per iteration (ms): 1065339.4 | throughput per GPU (TFLOP/s/GPU): 267.5 | MFU 27.05% | learning rate: 6.712553E-06 | global batch size: 1024 | lm loss: 3.847867E+00 | loss scale: 1.0 | grad norm: 581559296.000 | num zeros: 50149872.0 | params norm: 9733.412 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 17:09:49.684691 | finish at 2025-09-18 15:54:08
Number of parameters in transformer block in billions: 4.09
Number of parameters in embedding layers in billions: 0.38
Total number of parameters in billions: 4.47
Number of parameters in most loaded shard in billions: 2.2344
Activation memory footprint per transformer layer: 840.0 MB
Theoretical memory footprints: weight and optimizer=25570.57 MB, activation=100422.12 MB, total=125992.69 MB
[Rank 0] (after 1 iterations) memory (MB) | allocated: 12648.5341796875 | max allocated: 44991.19189453125 | reserved: 47224.0 | max reserved: 47224.0 | MEM: 48.54%
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.)
Variable._execution_engine.run_backward(
[2025-09-17 23:01:03] iteration 2/ 59 | consumed samples: 2048 | elapsed time per iteration (ms): 1016177.4 | throughput per GPU (TFLOP/s/GPU): 280.4 | MFU 28.36% | learning rate: 1.342511E-05 | global batch size: 1024 | lm loss: 3.936432E+00 | loss scale: 1.0 | grad norm: 2996916736.000 | num zeros: 49567872.0 | params norm: 9733.406 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 16:05:22.111929 | finish at 2025-09-18 15:06:25
[2025-09-17 23:17:48] iteration 3/ 59 | consumed samples: 3072 | elapsed time per iteration (ms): 1005115.4 | throughput per GPU (TFLOP/s/GPU): 283.5 | MFU 28.67% | learning rate: 1.999301E-05 | global batch size: 1024 | lm loss: 3.885512E+00 | loss scale: 1.0 | grad norm: 3171558400.000 | num zeros: 48767412.0 | params norm: 9733.392 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 15:38:06.463900 | finish at 2025-09-18 14:55:55
[2025-09-17 23:34:34] iteration 4/ 59 | consumed samples: 4096 | elapsed time per iteration (ms): 1005725.7 | throughput per GPU (TFLOP/s/GPU): 283.4 | MFU 28.65% | learning rate: 1.965217E-05 | global batch size: 1024 | lm loss: 3.872411E+00 | loss scale: 1.0 | grad norm: 28114848.000 | num zeros: 49829060.0 | params norm: 9733.373 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 15:21:54.915160 | finish at 2025-09-18 14:56:29
[2025-09-17 23:51:19] iteration 5/ 59 | consumed samples: 5120 | elapsed time per iteration (ms): 1004912.5 | throughput per GPU (TFLOP/s/GPU): 283.6 | MFU 28.67% | learning rate: 1.931133E-05 | global batch size: 1024 | lm loss: 3.757929E+00 | loss scale: 1.0 | grad norm: 1324076160.000 | num zeros: 49525848.0 | params norm: 9733.354 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 15:04:25.272742 | finish at 2025-09-18 14:55:44
Iter 5 reset timing-log-level.
[2025-09-18 00:08:00] iteration 6/ 59 | consumed samples: 6144 | elapsed time per iteration (ms): 1000789.1 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.897049E-05 | global batch size: 1024 | lm loss: 3.720438E+00 | loss scale: 1.0 | grad norm: 128727096.000 | num zeros: 48811584.0 | params norm: 9733.335 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 14:44:01.821778 | finish at 2025-09-18 14:52:01
[2025-09-18 00:24:41] iteration 7/ 59 | consumed samples: 7168 | elapsed time per iteration (ms): 1000967.1 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.862965E-05 | global batch size: 1024 | lm loss: 3.737079E+00 | loss scale: 1.0 | grad norm: 2022209152.000 | num zeros: 47870832.0 | params norm: 9733.317 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 14:27:30.288079 | finish at 2025-09-18 14:52:11
[2025-09-18 00:41:21] iteration 8/ 59 | consumed samples: 8192 | elapsed time per iteration (ms): 1000824.7 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.828882E-05 | global batch size: 1024 | lm loss: 3.762561E+00 | loss scale: 1.0 | grad norm: 5132039680.000 | num zeros: 49402816.0 | params norm: 9733.299 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 14:10:42.058004 | finish at 2025-09-18 14:52:03
[2025-09-18 00:58:02] iteration 9/ 59 | consumed samples: 9216 | elapsed time per iteration (ms): 1000899.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.794798E-05 | global batch size: 1024 | lm loss: 3.833430E+00 | loss scale: 1.0 | grad norm: 8436005888.000 | num zeros: 49748140.0 | params norm: 9733.281 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 13:54:04.963479 | finish at 2025-09-18 14:52:07
[2025-09-18 01:14:43] iteration 10/ 59 | consumed samples: 10240 | elapsed time per iteration (ms): 1000818.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.760714E-05 | global batch size: 1024 | lm loss: 3.975019E+00 | loss scale: 1.0 | grad norm: 13745026048.000 | num zeros: 48563464.0 | params norm: 9733.263 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 13:37:20.092904 | finish at 2025-09-18 14:52:03
[2025-09-18 01:31:24] iteration 11/ 59 | consumed samples: 11264 | elapsed time per iteration (ms): 1000845.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.726630E-05 | global batch size: 1024 | lm loss: 4.137877E+00 | loss scale: 1.0 | grad norm: 110488526848.000 | num zeros: 50133848.0 | params norm: 9733.247 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 13:20:40.568092 | finish at 2025-09-18 14:52:05
[2025-09-18 01:48:05] iteration 12/ 59 | consumed samples: 12288 | elapsed time per iteration (ms): 1000906.4 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.692546E-05 | global batch size: 1024 | lm loss: 4.174570E+00 | loss scale: 1.0 | grad norm: 213260451840.000 | num zeros: 49354312.0 | params norm: 9733.230 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 13:04:02.599566 | finish at 2025-09-18 14:52:07
[2025-09-18 02:04:46] iteration 13/ 59 | consumed samples: 13312 | elapsed time per iteration (ms): 1000826.5 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.658462E-05 | global batch size: 1024 | lm loss: 4.318902E+00 | loss scale: 1.0 | grad norm: 29837592576.000 | num zeros: 49300288.0 | params norm: 9733.213 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 12:47:18.019107 | finish at 2025-09-18 14:52:04
[2025-09-18 02:21:26] iteration 14/ 59 | consumed samples: 14336 | elapsed time per iteration (ms): 1000652.2 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.80% | learning rate: 1.624378E-05 | global batch size: 1024 | lm loss: 4.142406E+00 | loss scale: 1.0 | grad norm: 5777787904.000 | num zeros: 47241792.0 | params norm: 9733.197 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 12:30:29.348720 | finish at 2025-09-18 14:51:56
[2025-09-18 02:38:07] iteration 15/ 59 | consumed samples: 15360 | elapsed time per iteration (ms): 1000613.6 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.80% | learning rate: 1.590294E-05 | global batch size: 1024 | lm loss: 4.086899E+00 | loss scale: 1.0 | grad norm: 48051974144.000 | num zeros: 49042056.0 | params norm: 9733.181 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 12:13:46.997184 | finish at 2025-09-18 14:51:54
[2025-09-18 02:54:48] iteration 16/ 59 | consumed samples: 16384 | elapsed time per iteration (ms): 1000891.7 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.556211E-05 | global batch size: 1024 | lm loss: 3.986907E+00 | loss scale: 1.0 | grad norm: 1339338496.000 | num zeros: 49090020.0 | params norm: 9733.166 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 11:57:18.344998 | finish at 2025-09-18 14:52:06
[2025-09-18 03:11:29] iteration 17/ 59 | consumed samples: 17408 | elapsed time per iteration (ms): 1000751.6 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.522127E-05 | global batch size: 1024 | lm loss: 3.926814E+00 | loss scale: 1.0 | grad norm: 751090048.000 | num zeros: 49013412.0 | params norm: 9733.150 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 11:40:31.568773 | finish at 2025-09-18 14:52:00
[2025-09-18 03:28:09] iteration 18/ 59 | consumed samples: 18432 | elapsed time per iteration (ms): 1000766.6 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.488043E-05 | global batch size: 1024 | lm loss: 3.843276E+00 | loss scale: 1.0 | grad norm: 4079015936.000 | num zeros: 48842440.0 | params norm: 9733.136 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 11:23:51.429882 | finish at 2025-09-18 14:52:01
[2025-09-18 03:44:50] iteration 19/ 59 | consumed samples: 19456 | elapsed time per iteration (ms): 1000849.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.453959E-05 | global batch size: 1024 | lm loss: 3.819354E+00 | loss scale: 1.0 | grad norm: 663727040.000 | num zeros: 48992272.0 | params norm: 9733.121 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 11:07:13.967733 | finish at 2025-09-18 14:52:04
[2025-09-18 04:01:31] iteration 20/ 59 | consumed samples: 20480 | elapsed time per iteration (ms): 1000851.7 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.419875E-05 | global batch size: 1024 | lm loss: 3.898749E+00 | loss scale: 1.0 | grad norm: 177677072.000 | num zeros: 49040544.0 | params norm: 9733.107 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 10:50:33.214592 | finish at 2025-09-18 14:52:04
[2025-09-18 04:18:12] iteration 21/ 59 | consumed samples: 21504 | elapsed time per iteration (ms): 1000741.4 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.385791E-05 | global batch size: 1024 | lm loss: 4.066440E+00 | loss scale: 1.0 | grad norm: 102174968.000 | num zeros: 49475984.0 | params norm: 9733.093 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 10:33:48.171461 | finish at 2025-09-18 14:52:00
[2025-09-18 04:34:53] iteration 22/ 59 | consumed samples: 22528 | elapsed time per iteration (ms): 1000731.2 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.351707E-05 | global batch size: 1024 | lm loss: 4.138809E+00 | loss scale: 1.0 | grad norm: 1939881984.000 | num zeros: 49917468.0 | params norm: 9733.080 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 10:17:07.053182 | finish at 2025-09-18 14:52:00
[2025-09-18 04:51:33] iteration 23/ 59 | consumed samples: 23552 | elapsed time per iteration (ms): 1000630.4 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.80% | learning rate: 1.317623E-05 | global batch size: 1024 | lm loss: 4.354475E+00 | loss scale: 1.0 | grad norm: 1467312896.000 | num zeros: 49097692.0 | params norm: 9733.067 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 10:00:22.694501 | finish at 2025-09-18 14:51:56
[2025-09-18 05:08:14] iteration 24/ 59 | consumed samples: 24576 | elapsed time per iteration (ms): 1000795.8 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.283539E-05 | global batch size: 1024 | lm loss: 4.601981E+00 | loss scale: 1.0 | grad norm: 23645913088.000 | num zeros: 49239816.0 | params norm: 9733.054 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 9:43:47.853149 | finish at 2025-09-18 14:52:02
[2025-09-18 05:24:55] iteration 25/ 59 | consumed samples: 25600 | elapsed time per iteration (ms): 1000910.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.249456E-05 | global batch size: 1024 | lm loss: 4.786105E+00 | loss scale: 1.0 | grad norm: 12307973120.000 | num zeros: 49598280.0 | params norm: 9733.041 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 9:27:10.949439 | finish at 2025-09-18 14:52:06
[2025-09-18 05:41:36] iteration 26/ 59 | consumed samples: 26624 | elapsed time per iteration (ms): 1000848.6 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.215372E-05 | global batch size: 1024 | lm loss: 5.074014E+00 | loss scale: 1.0 | grad norm: 26155917312.000 | num zeros: 49150844.0 | params norm: 9733.029 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 9:10:28.003309 | finish at 2025-09-18 14:52:04
[2025-09-18 05:58:17] iteration 27/ 59 | consumed samples: 27648 | elapsed time per iteration (ms): 1000920.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.181288E-05 | global batch size: 1024 | lm loss: 5.265514E+00 | loss scale: 1.0 | grad norm: 124921552896.000 | num zeros: 49020472.0 | params norm: 9733.018 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 8:53:49.449287 | finish at 2025-09-18 14:52:06
[2025-09-18 06:14:58] iteration 28/ 59 | consumed samples: 28672 | elapsed time per iteration (ms): 1000958.1 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.147204E-05 | global batch size: 1024 | lm loss: 5.399610E+00 | loss scale: 1.0 | grad norm: 7730917888.000 | num zeros: 49531044.0 | params norm: 9733.006 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 8:37:09.699772 | finish at 2025-09-18 14:52:07
[2025-09-18 06:31:38] iteration 29/ 59 | consumed samples: 29696 | elapsed time per iteration (ms): 1000807.7 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.113120E-05 | global batch size: 1024 | lm loss: 5.464283E+00 | loss scale: 1.0 | grad norm: 1776406757376.000 | num zeros: 49643804.0 | params norm: 9732.995 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 8:20:24.230998 | finish at 2025-09-18 14:52:03
saving checkpoint at iteration 29 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 in torch format
successfully saved checkpoint from iteration 29 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 [ t 1/2, p 1/1 ]
(min, max) time across ranks (ms):
save-checkpoint ................................: (51941.34, 51941.37)
[2025-09-18 06:49:11] iteration 30/ 59 | consumed samples: 30720 | elapsed time per iteration (ms): 1001022.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.079036E-05 | global batch size: 1024 | lm loss: 5.466630E+00 | loss scale: 1.0 | grad norm: 201713991680.000 | num zeros: 48573168.0 | params norm: 9732.984 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 8:03:49.644004 | finish at 2025-09-18 14:53:01
[2025-09-18 07:05:52] iteration 31/ 59 | consumed samples: 31744 | elapsed time per iteration (ms): 1000966.6 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.044952E-05 | global batch size: 1024 | lm loss: 5.530435E+00 | loss scale: 1.0 | grad norm: 16827718656.000 | num zeros: 49106456.0 | params norm: 9732.973 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 7:47:07.063577 | finish at 2025-09-18 14:52:59
[2025-09-18 07:22:33] iteration 32/ 59 | consumed samples: 32768 | elapsed time per iteration (ms): 1001076.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 1.010868E-05 | global batch size: 1024 | lm loss: 5.704006E+00 | loss scale: 1.0 | grad norm: 279325507584.000 | num zeros: 48570972.0 | params norm: 9732.963 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 7:30:29.056911 | finish at 2025-09-18 14:53:02
[2025-09-18 07:39:14] iteration 33/ 59 | consumed samples: 33792 | elapsed time per iteration (ms): 1000973.1 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 9.767845E-06 | global batch size: 1024 | lm loss: 5.747099E+00 | loss scale: 1.0 | grad norm: 31260119040.000 | num zeros: 49229308.0 | params norm: 9732.954 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 7:13:45.300946 | finish at 2025-09-18 14:53:00
[2025-09-18 07:55:55] iteration 34/ 59 | consumed samples: 34816 | elapsed time per iteration (ms): 1000882.6 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 9.427005E-06 | global batch size: 1024 | lm loss: 5.803835E+00 | loss scale: 1.0 | grad norm: 640638255104.000 | num zeros: 49311708.0 | params norm: 9732.944 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 6:57:02.066164 | finish at 2025-09-18 14:52:57
[2025-09-18 08:12:36] iteration 35/ 59 | consumed samples: 35840 | elapsed time per iteration (ms): 1000906.9 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 9.086167E-06 | global batch size: 1024 | lm loss: 5.908048E+00 | loss scale: 1.0 | grad norm: 101146836992.000 | num zeros: 48868296.0 | params norm: 9732.935 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 6:40:21.766422 | finish at 2025-09-18 14:52:58
[2025-09-18 08:29:17] iteration 36/ 59 | consumed samples: 36864 | elapsed time per iteration (ms): 1000991.8 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 8.745328E-06 | global batch size: 1024 | lm loss: 6.095221E+00 | loss scale: 1.0 | grad norm: 461121159168.000 | num zeros: 48946824.0 | params norm: 9732.926 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 6:23:42.811659 | finish at 2025-09-18 14:53:00
[2025-09-18 08:45:58] iteration 37/ 59 | consumed samples: 37888 | elapsed time per iteration (ms): 1000857.1 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 8.404490E-06 | global batch size: 1024 | lm loss: 5.988223E+00 | loss scale: 1.0 | grad norm: 50295472128.000 | num zeros: 49686208.0 | params norm: 9732.918 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 6:06:58.856651 | finish at 2025-09-18 14:52:57
[2025-09-18 09:02:39] iteration 38/ 59 | consumed samples: 38912 | elapsed time per iteration (ms): 1001077.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 8.063650E-06 | global batch size: 1024 | lm loss: 5.932316E+00 | loss scale: 1.0 | grad norm: 465845452800.000 | num zeros: 49469216.0 | params norm: 9732.909 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 5:50:22.621344 | finish at 2025-09-18 14:53:02
[2025-09-18 09:19:20] iteration 39/ 59 | consumed samples: 39936 | elapsed time per iteration (ms): 1000904.9 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 7.722811E-06 | global batch size: 1024 | lm loss: 5.991258E+00 | loss scale: 1.0 | grad norm: 100241760256.000 | num zeros: 49582496.0 | params norm: 9732.901 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 5:33:38.097820 | finish at 2025-09-18 14:52:58
[2025-09-18 09:36:01] iteration 40/ 59 | consumed samples: 40960 | elapsed time per iteration (ms): 1001159.5 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 7.381972E-06 | global batch size: 1024 | lm loss: 5.912158E+00 | loss scale: 1.0 | grad norm: 39360831488.000 | num zeros: 49063880.0 | params norm: 9732.894 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 5:17:02.031186 | finish at 2025-09-18 14:53:03
[2025-09-18 09:52:42] iteration 41/ 59 | consumed samples: 41984 | elapsed time per iteration (ms): 1001076.8 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 7.041134E-06 | global batch size: 1024 | lm loss: 5.951908E+00 | loss scale: 1.0 | grad norm: 159719407616.000 | num zeros: 47581900.0 | params norm: 9732.887 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 5:00:19.381758 | finish at 2025-09-18 14:53:02
[2025-09-18 10:09:23] iteration 42/ 59 | consumed samples: 43008 | elapsed time per iteration (ms): 1000910.9 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 6.700295E-06 | global batch size: 1024 | lm loss: 5.992692E+00 | loss scale: 1.0 | grad norm: 198826213376.000 | num zeros: 49029128.0 | params norm: 9732.880 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 4:43:35.484690 | finish at 2025-09-18 14:52:59
[2025-09-18 10:26:04] iteration 43/ 59 | consumed samples: 44032 | elapsed time per iteration (ms): 1001109.6 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 6.359456E-06 | global batch size: 1024 | lm loss: 6.063922E+00 | loss scale: 1.0 | grad norm: 81292369920.000 | num zeros: 48821880.0 | params norm: 9732.874 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 4:26:57.753551 | finish at 2025-09-18 14:53:02
[2025-09-18 10:42:45] iteration 44/ 59 | consumed samples: 45056 | elapsed time per iteration (ms): 1001102.0 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 6.018617E-06 | global batch size: 1024 | lm loss: 6.145180E+00 | loss scale: 1.0 | grad norm: 217426198528.000 | num zeros: 49321256.0 | params norm: 9732.868 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 4:10:16.529392 | finish at 2025-09-18 14:53:02
[2025-09-18 10:59:26] iteration 45/ 59 | consumed samples: 46080 | elapsed time per iteration (ms): 1000832.6 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 5.677779E-06 | global batch size: 1024 | lm loss: 6.157466E+00 | loss scale: 1.0 | grad norm: 96730759168.000 | num zeros: 49428652.0 | params norm: 9732.861 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 3:53:31.656602 | finish at 2025-09-18 14:52:58
[2025-09-18 11:16:07] iteration 46/ 59 | consumed samples: 47104 | elapsed time per iteration (ms): 1001193.6 | throughput per GPU (TFLOP/s/GPU): 284.6 | MFU 28.78% | learning rate: 5.336939E-06 | global batch size: 1024 | lm loss: 6.188946E+00 | loss scale: 1.0 | grad norm: 666173046784.000 | num zeros: 48664652.0 | params norm: 9732.856 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 3:36:55.516861 | finish at 2025-09-18 14:53:03
[2025-09-18 11:32:48] iteration 47/ 59 | consumed samples: 48128 | elapsed time per iteration (ms): 1001008.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 4.996101E-06 | global batch size: 1024 | lm loss: 6.225647E+00 | loss scale: 1.0 | grad norm: 185815629824.000 | num zeros: 49034972.0 | params norm: 9732.851 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 3:20:12.100107 | finish at 2025-09-18 14:53:01
[2025-09-18 11:49:30] iteration 48/ 59 | consumed samples: 49152 | elapsed time per iteration (ms): 1001083.1 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 4.655262E-06 | global batch size: 1024 | lm loss: 6.273437E+00 | loss scale: 1.0 | grad norm: 534045065216.000 | num zeros: 48900832.0 | params norm: 9732.846 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 3:03:31.914597 | finish at 2025-09-18 14:53:01
[2025-09-18 12:06:10] iteration 49/ 59 | consumed samples: 50176 | elapsed time per iteration (ms): 1000901.7 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 4.314423E-06 | global batch size: 1024 | lm loss: 6.338528E+00 | loss scale: 1.0 | grad norm: 899446276096.000 | num zeros: 48965872.0 | params norm: 9732.841 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 2:46:49.016547 | finish at 2025-09-18 14:52:59
[2025-09-18 12:22:51] iteration 50/ 59 | consumed samples: 51200 | elapsed time per iteration (ms): 1000783.5 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 3.973584E-06 | global batch size: 1024 | lm loss: 6.234092E+00 | loss scale: 1.0 | grad norm: 199767703552.000 | num zeros: 49578360.0 | params norm: 9732.838 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 2:30:07.051476 | finish at 2025-09-18 14:52:58
[2025-09-18 12:39:32] iteration 51/ 59 | consumed samples: 52224 | elapsed time per iteration (ms): 1000899.7 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 3.632745E-06 | global batch size: 1024 | lm loss: 6.373925E+00 | loss scale: 1.0 | grad norm: 253297819648.000 | num zeros: 49478376.0 | params norm: 9732.834 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 2:13:27.197994 | finish at 2025-09-18 14:52:59
[2025-09-18 12:56:13] iteration 52/ 59 | consumed samples: 53248 | elapsed time per iteration (ms): 1001001.0 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 3.291906E-06 | global batch size: 1024 | lm loss: 6.419476E+00 | loss scale: 1.0 | grad norm: 113803444224.000 | num zeros: 49280460.0 | params norm: 9732.830 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 1:56:47.007308 | finish at 2025-09-18 14:53:00
[2025-09-18 13:12:54] iteration 53/ 59 | consumed samples: 54272 | elapsed time per iteration (ms): 1001045.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 2.951067E-06 | global batch size: 1024 | lm loss: 6.387888E+00 | loss scale: 1.0 | grad norm: 255267913728.000 | num zeros: 47073088.0 | params norm: 9732.827 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 1:40:06.271739 | finish at 2025-09-18 14:53:00
[2025-09-18 13:29:35] iteration 54/ 59 | consumed samples: 55296 | elapsed time per iteration (ms): 1001091.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 2.610229E-06 | global batch size: 1024 | lm loss: 6.394367E+00 | loss scale: 1.0 | grad norm: 131194019840.000 | num zeros: 49299976.0 | params norm: 9732.824 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 1:23:25.456586 | finish at 2025-09-18 14:53:01
[2025-09-18 13:46:16] iteration 55/ 59 | consumed samples: 56320 | elapsed time per iteration (ms): 1000805.8 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 2.269390E-06 | global batch size: 1024 | lm loss: 6.428160E+00 | loss scale: 1.0 | grad norm: 526357069824.000 | num zeros: 49432300.0 | params norm: 9732.822 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 1:06:43.223324 | finish at 2025-09-18 14:52:59
[2025-09-18 14:02:57] iteration 56/ 59 | consumed samples: 57344 | elapsed time per iteration (ms): 1000877.8 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.928551E-06 | global batch size: 1024 | lm loss: 6.425514E+00 | loss scale: 1.0 | grad norm: 80378552320.000 | num zeros: 49106088.0 | params norm: 9732.819 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 0:50:02.633369 | finish at 2025-09-18 14:53:00
[2025-09-18 14:19:38] iteration 57/ 59 | consumed samples: 58368 | elapsed time per iteration (ms): 1000950.9 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.587712E-06 | global batch size: 1024 | lm loss: 6.411983E+00 | loss scale: 1.0 | grad norm: 601219596288.000 | num zeros: 49628192.0 | params norm: 9732.817 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 0:33:21.901740 | finish at 2025-09-18 14:53:00
[2025-09-18 14:36:19] iteration 58/ 59 | consumed samples: 59392 | elapsed time per iteration (ms): 1000958.0 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.246873E-06 | global batch size: 1024 | lm loss: 6.455757E+00 | loss scale: 1.0 | grad norm: 586592092160.000 | num zeros: 48973652.0 | params norm: 9732.816 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 0:16:40.958041 | finish at 2025-09-18 14:53:00
saving checkpoint at iteration 58 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 in torch format
successfully saved checkpoint from iteration 58 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 [ t 1/2, p 1/1 ]
(min, max) time across ranks (ms):
save-checkpoint ................................: (51718.88, 51718.90)
[2025-09-18 14:53:51] iteration 59/ 59 | consumed samples: 60416 | elapsed time per iteration (ms): 1000781.5 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 9.060344E-07 | global batch size: 1024 | lm loss: 6.509560E+00 | loss scale: 1.0 | grad norm: 141197361152.000 | num zeros: 49217292.0 | params norm: 9732.815 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 0:00:00 | finish at 2025-09-18 14:53:51
[after training is done] datetime: 2025-09-18 14:53:51
saving checkpoint at iteration 59 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 in torch format
successfully saved checkpoint from iteration 59 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 [ t 1/2, p 1/1 ]
[rank1]:[W918 14:54:44.101205233 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank3]:[W918 14:54:44.170021956 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank2]:[W918 14:54:44.807181589 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank0]:[W918 14:54:45.828822024 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank4]:[W918 14:54:45.679057092 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank6]:[W918 14:54:45.727992159 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank7]:[W918 14:54:45.747321323 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank5]:[W918 14:54:46.100354521 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())