| using world size: 8, data-parallel size: 2, context-parallel size: 2, hierarchical context-parallel sizes: Nonetensor-model-parallel size: 2, encoder-tensor-model-parallel size: 2, pipeline-model-parallel size: 1, encoder-pipeline-model-parallel size: 0 |
| WARNING: overriding default arguments for tokenizer_type:GPT2BPETokenizer with tokenizer_type:HuggingFaceTokenizer |
| Number of virtual stages per pipeline stage: None |
| accumulate and all-reduce gradients in fp32 for bfloat16 data type. |
| using torch.bfloat16 for parameters ... |
| ------------------------ arguments ------------------------ |
| account_for_embedding_in_pipeline_split ......... False |
| account_for_loss_in_pipeline_split .............. False |
| accumulate_allreduce_grads_in_fp32 .............. True |
| adam_beta1 ...................................... 0.9 |
| adam_beta2 ...................................... 0.999 |
| adam_eps ........................................ 1e-08 |
| add_bias_linear ................................. False |
| add_position_embedding .......................... False |
| add_qkv_bias .................................... True |
| adlr_autoresume ................................. False |
| adlr_autoresume_interval ........................ 1000 |
| align_grad_reduce ............................... True |
| align_param_gather .............................. False |
| app_tag_run_name ................................ None |
| app_tag_run_version ............................. 0.0.0 |
| apply_layernorm_1p .............................. False |
| apply_query_key_layer_scaling ................... False |
| apply_residual_connection_post_layernorm ........ False |
| apply_rope_fusion ............................... True |
| async_save ...................................... None |
| async_tensor_model_parallel_allreduce ........... True |
| attention_backend ............................... AttnBackend.auto |
| attention_dropout ............................... 0.0 |
| attention_softmax_in_fp32 ....................... False |
| attn_output_gate ................................ None |
| attn_token_shift ................................ None |
| auto_detect_ckpt_format ......................... False |
| barrier_with_L1_time ............................ True |
| bert_binary_head ................................ True |
| bert_embedder_type .............................. megatron |
| bert_load ....................................... None |
| bf16 ............................................ True |
| bias_dropout_fusion ............................. True |
| bias_gelu_fusion ................................ False |
| bias_swiglu_fusion .............................. True |
| biencoder_projection_dim ........................ 0 |
| biencoder_shared_query_context_model ............ False |
| block_data_path ................................. None |
| calc_ft_timeouts ................................ False |
| calculate_per_token_loss ........................ False |
| check_for_large_grads ........................... False |
| check_for_nan_in_loss_and_grad .................. True |
| check_for_spiky_loss ............................ False |
| check_weight_hash_across_dp_replicas_interval ... None |
| ckpt_assume_constant_structure .................. False |
| ckpt_convert_format ............................. None |
| ckpt_convert_save ............................... None |
| ckpt_convert_update_legacy_dist_opt_format ...... False |
| ckpt_format ..................................... torch |
| ckpt_fully_parallel_load ........................ False |
| ckpt_fully_parallel_save ........................ True |
| ckpt_fully_parallel_save_deprecated ............. False |
| ckpt_step ....................................... None |
| classes_fraction ................................ 1.0 |
| clip_grad ....................................... 0.5 |
| clone_scatter_output_in_embedding ............... True |
| config_logger_dir ............................... |
| consumed_train_samples .......................... 0 |
| consumed_valid_samples .......................... 0 |
| context_parallel_size ........................... 2 |
| cp_comm_type .................................... ['p2p'] |
| create_attention_mask_in_dataloader ............. False |
| cross_entropy_fusion_impl ....................... native |
| cross_entropy_loss_fusion ....................... False |
| cuda_graph_scope ................................ full |
| cuda_graph_warmup_steps ......................... 3 |
| data_args_path .................................. None |
| data_cache_path ................................. /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache |
| data_parallel_random_init ....................... False |
| data_parallel_sharding_strategy ................. no_shard |
| data_parallel_size .............................. 2 |
| data_path ....................................... ['/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache/datasets/huggingface/Teaven/combine_2B_0908/binidx/yulan_mini'] |
| data_per_class_fraction ......................... 1.0 |
| data_sharding ................................... True |
| dataloader_type ................................. single |
| ddp_average_in_collective ....................... False |
| ddp_bucket_size ................................. None |
| ddp_num_buckets ................................. None |
| ddp_pad_buckets_for_high_nccl_busbw ............. False |
| decoder_first_pipeline_num_layers ............... None |
| decoder_last_pipeline_num_layers ................ None |
| decoder_num_layers .............................. None |
| decoder_seq_length .............................. None |
| decoupled_lr .................................... None |
| decoupled_min_lr ................................ None |
| decrease_batch_size_if_needed ................... False |
| defer_embedding_wgrad_compute ................... False |
| deprecated_use_mcore_models ..................... True |
| deterministic_mode .............................. False |
| dino_bottleneck_size ............................ 256 |
| dino_freeze_last_layer .......................... 1 |
| dino_head_hidden_size ........................... 2048 |
| dino_local_crops_number ......................... 10 |
| dino_local_img_size ............................. 96 |
| dino_norm_last_layer ............................ False |
| dino_teacher_temp ............................... 0.07 |
| dino_warmup_teacher_temp ........................ 0.04 |
| dino_warmup_teacher_temp_epochs ................. 30 |
| disable_bf16_reduced_precision_matmul ........... False |
| disable_mamba_mem_eff_path ...................... False |
| disable_straggler_on_startup .................... False |
| dist_ckpt_format_deprecated ..................... None |
| dist_ckpt_strictness ............................ assume_ok_unexpected |
| distribute_saved_activations .................... False |
| distributed_backend ............................. nccl |
| distributed_timeout_minutes ..................... 10 |
| emb_deviation_loss_coeff ........................ 0 |
| emb_deviation_type .............................. None |
| embedding_path .................................. None |
| empty_unused_memory_level ....................... 0 |
| enable_cuda_graph ............................... False |
| enable_ft_package ............................... False |
| enable_gloo_process_groups ...................... True |
| enable_msc ...................................... True |
| enable_one_logger ............................... True |
| encoder_num_layers .............................. 112 |
| encoder_pipeline_model_parallel_size ............ 0 |
| encoder_seq_length .............................. 32768 |
| encoder_tensor_model_parallel_size .............. 2 |
| end_weight_decay ................................ 0.1 |
| eod_mask_loss ................................... False |
| error_injection_rate ............................ 0 |
| error_injection_type ............................ transient_error |
| eval_interval ................................... 1000 |
| eval_iters ...................................... 10 |
| evidence_data_path .............................. None |
| exit_duration_in_mins ........................... None |
| exit_interval ................................... None |
| exit_on_missing_checkpoint ...................... False |
| exit_signal_handler ............................. False |
| exp_avg_dtype ................................... torch.float32 |
| exp_avg_sq_dtype ................................ torch.float32 |
| expert_model_parallel_size ...................... 1 |
| expert_tensor_parallel_size ..................... 2 |
| external_cuda_graph ............................. False |
| ffn_hidden_size ................................. 4800 |
| ffn_token_shift ................................. None |
| finetune ........................................ False |
| first_last_layers_bf16 .......................... False |
| flash_decode .................................... False |
| fp16 ............................................ False |
| fp16_lm_cross_entropy ........................... False |
| fp32_residual_connection ........................ False |
| fp8 ............................................. None |
| fp8_amax_compute_algo ........................... most_recent |
| fp8_amax_history_len ............................ 1 |
| fp8_interval .................................... 1 |
| fp8_margin ...................................... 0 |
| fp8_param_gather ................................ False |
| fp8_recipe ...................................... delayed |
| fp8_wgrad ....................................... True |
| freeze_non_mamba ................................ False |
| geglu ........................................... False |
| global_batch_size ............................... 1024 |
| grad_reduce_in_bf16 ............................. False |
| gradient_accumulation_fusion .................... True |
| gradient_reduce_div_fusion ...................... True |
| group_query_attention ........................... True |
| head_lr_mult .................................... 1.0 |
| heterogeneous_layers_config_encoded_json ........ None |
| heterogeneous_layers_config_path ................ None |
| hidden_dropout .................................. 0.0 |
| hidden_size ..................................... 1920 |
| hierarchical_context_parallel_sizes ............. None |
| hybrid_attention_ratio .......................... 0.0625 |
| hybrid_mlp_ratio ................................ 0.5 |
| hybrid_override_pattern ......................... *-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M- |
| hysteresis ...................................... 2 |
| ict_head_size ................................... None |
| ict_load ........................................ None |
| img_h ........................................... 224 |
| img_w ........................................... 224 |
| increase_log_level_interval ..................... 1000 |
| increase_log_level_iters ........................ 5 |
| indexer_batch_size .............................. 128 |
| indexer_log_interval ............................ 1000 |
| inference_batch_times_seqlen_threshold .......... -1 |
| inference_dynamic_batching ...................... False |
| inference_dynamic_batching_buffer_guaranteed_fraction 0.2 |
| inference_dynamic_batching_buffer_overflow_factor None |
| inference_dynamic_batching_buffer_size_gb ....... 40.0 |
| inference_dynamic_batching_chunk_size ........... 256 |
| inference_dynamic_batching_max_requests_override None |
| inference_dynamic_batching_max_tokens_override .. None |
| inference_max_batch_size ........................ 8 |
| inference_max_seq_length ........................ 2560 |
| inference_rng_tracker ........................... False |
| init_method_std ................................. 0.02 |
| init_method_xavier_uniform ...................... False |
| init_model_with_meta_device ..................... False |
| initial_loss_scale .............................. 4294967296 |
| is_hybrid_model ................................. False |
| iter_per_epoch .................................. 1250 |
| iterations_to_skip .............................. [] |
| keep_fp8_transpose_cache_when_using_custom_fsdp . False |
| kv_channels ..................................... 64 |
| kv_lora_rank .................................... 32 |
| lazy_mpu_init ................................... None |
| load ............................................ /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/RADLADS-paper/out/L56-D1920-qwen_mamba2_qwen2-e1-i1920-s320-hd64-gn6-A0-S512--step1-dclm10b/rwkv-394-hf-A7-0_8_16_24_32_40_48/megatron-pp1-tp2 |
| local_rank ...................................... 0 |
| log_interval .................................... 1 |
| log_layer_hidden_states ......................... [] |
| log_loss_scale_to_tensorboard ................... True |
| log_memory_to_tensorboard ....................... True |
| log_num_zeros_in_grad ........................... False |
| log_params_norm ................................. True |
| log_progress .................................... False |
| log_straggler ................................... False |
| log_throughput .................................. True |
| log_timers_to_tensorboard ....................... True |
| log_validation_ppl_to_tensorboard ............... False |
| log_world_size_to_tensorboard ................... False |
| logging_level ................................... None |
| loss_scale ...................................... None |
| loss_scale_window ............................... 1000 |
| lr .............................................. 2e-05 |
| lr_decay_iters .................................. None |
| lr_decay_samples ................................ 61035 |
| lr_decay_style .................................. linear |
| lr_warmup_fraction .............................. None |
| lr_warmup_init .................................. 0.0 |
| lr_warmup_iters ................................. 0 |
| lr_warmup_samples ............................... 3051 |
| lr_wsd_decay_iters .............................. None |
| lr_wsd_decay_samples ............................ None |
| lr_wsd_decay_style .............................. exponential |
| main_grads_dtype ................................ torch.float32 |
| main_params_dtype ............................... torch.float32 |
| make_vocab_size_divisible_by .................... 128 |
| mamba_expand .................................... 1 |
| mamba_head_dim .................................. 64 |
| mamba_num_groups ................................ 6 |
| mamba_num_heads ................................. None |
| mamba_state_dim ................................. 320 |
| manual_gc ....................................... False |
| manual_gc_eval .................................. True |
| manual_gc_interval .............................. 0 |
| mask_factor ..................................... 1.0 |
| mask_prob ....................................... 0.15 |
| mask_type ....................................... random |
| masked_softmax_fusion ........................... False |
| max_position_embeddings ......................... 32768 |
| max_tokens_to_oom ............................... 12000 |
| memory_snapshot_path ............................ snapshot.pickle |
| merge_file ...................................... None |
| micro_batch_size ................................ 1 |
| microbatch_group_size_per_vp_stage .............. None |
| mid_level_dataset_surplus ....................... 0.005 |
| min_loss_scale .................................. 1.0 |
| min_lr .......................................... 7e-07 |
| mlp_chunks_for_prefill .......................... 1 |
| mmap_bin_files .................................. True |
| mock_data ....................................... False |
| moe_aux_loss_coeff .............................. 0.0 |
| moe_enable_deepep ............................... False |
| moe_expert_capacity_factor ...................... None |
| moe_extended_tp ................................. False |
| moe_ffn_hidden_size ............................. None |
| moe_grouped_gemm ................................ False |
| moe_input_jitter_eps ............................ None |
| moe_layer_freq .................................. 1 |
| moe_layer_recompute ............................. False |
| moe_pad_expert_input_to_capacity ................ False |
| moe_per_layer_logging ........................... False |
| moe_permute_fusion .............................. False |
| moe_router_bias_update_method ................... sign |
| moe_router_bias_update_rate ..................... 0.001 |
| moe_router_dtype ................................ None |
| moe_router_enable_expert_bias ................... False |
| moe_router_group_topk ........................... None |
| moe_router_load_balancing_type .................. aux_loss |
| moe_router_num_groups ........................... None |
| moe_router_pre_softmax .......................... False |
| moe_router_score_function ....................... softmax |
| moe_router_topk ................................. 2 |
| moe_router_topk_scaling_factor .................. None |
| moe_shared_expert_intermediate_size ............. None |
| moe_shared_expert_overlap ....................... False |
| moe_token_dispatcher_type ....................... allgather |
| moe_token_drop_policy ........................... probs |
| moe_use_legacy_grouped_gemm ..................... False |
| moe_use_upcycling ............................... False |
| moe_z_loss_coeff ................................ None |
| mrope_section ................................... None |
| mscale .......................................... 1.0 |
| mscale_all_dim .................................. 1.0 |
| mtp_loss_scaling_factor ......................... 0.1 |
| mtp_num_layers .................................. None |
| multi_latent_attention .......................... False |
| muon_matched_adamw_rms .......................... 0.2 |
| muon_momentum ................................... 0.95 |
| muon_nesterov ................................... True |
| muon_ns_steps ................................... 5 |
| nccl_communicator_config_path ................... None |
| no_load_optim ................................... True |
| no_load_rng ..................................... True |
| no_persist_layer_norm ........................... False |
| no_save_optim ................................... None |
| no_save_rng ..................................... None |
| no_save_step_one ................................ True |
| non_persistent_ckpt_type ........................ None |
| non_persistent_global_ckpt_dir .................. None |
| non_persistent_local_ckpt_algo .................. fully_parallel |
| non_persistent_local_ckpt_dir ................... None |
| non_persistent_save_interval .................... None |
| norm_epsilon .................................... 1e-05 |
| normalization ................................... RMSNorm |
| num_attention_heads ............................. 30 |
| num_channels .................................... 3 |
| num_classes ..................................... 1000 |
| num_dataset_builder_threads ..................... 1 |
| num_distributed_optimizer_instances ............. 1 |
| num_experts ..................................... None |
| num_layers ...................................... 112 |
| num_layers_at_end_in_bf16 ....................... 1 |
| num_layers_at_start_in_bf16 ..................... 1 |
| num_layers_per_virtual_pipeline_stage ........... None |
| num_query_groups ................................ 6 |
| num_virtual_stages_per_pipeline_rank ............ None |
| num_workers ..................................... 2 |
| object_storage_cache_path ....................... None |
| one_logger_async ................................ False |
| one_logger_project .............................. megatron-lm |
| one_logger_run_name ............................. None |
| onnx_safe ....................................... None |
| openai_gelu ..................................... False |
| optimizer ....................................... adam |
| optimizer_cpu_offload ........................... False |
| optimizer_offload_fraction ...................... 1.0 |
| output_bert_embeddings .......................... False |
| overlap_cpu_optimizer_d2h_h2d ................... False |
| overlap_grad_reduce ............................. True |
| overlap_p2p_comm ................................ False |
| overlap_p2p_comm_warmup_flush ................... False |
| overlap_param_gather ............................ True |
| overlap_param_gather_with_optimizer_step ........ False |
| override_opt_param_scheduler .................... False |
| params_dtype .................................... torch.bfloat16 |
| patch_dim ....................................... 16 |
| per_split_data_args_path ........................ None |
| perform_initialization .......................... True |
| pin_cpu_grads ................................... True |
| pin_cpu_params .................................. True |
| pipeline_model_parallel_comm_backend ............ None |
| pipeline_model_parallel_size .................... 1 |
| pipeline_model_parallel_split_rank .............. None |
| position_embedding_type ......................... rope |
| pretrained_checkpoint ........................... None |
| profile ......................................... False |
| profile_ranks ................................... [0] |
| profile_step_end ................................ 12 |
| profile_step_start .............................. 10 |
| q_lora_rank ..................................... None |
| qk_head_dim ..................................... 128 |
| qk_l2_norm ...................................... False |
| qk_layernorm .................................... False |
| qk_pos_emb_head_dim ............................. 64 |
| query_in_block_prob ............................. 0.1 |
| rampup_batch_size ............................... None |
| rank ............................................ 0 |
| recompute_granularity ........................... selective |
| recompute_method ................................ None |
| recompute_modules ............................... None |
| recompute_num_layers ............................ None |
| record_memory_history ........................... False |
| relative_attention_max_distance ................. 128 |
| relative_attention_num_buckets .................. 32 |
| replication ..................................... False |
| replication_factor .............................. 2 |
| replication_jump ................................ None |
| rerun_mode ...................................... disabled |
| reset_attention_mask ............................ False |
| reset_position_ids .............................. False |
| result_rejected_tracker_filename ................ None |
| retriever_report_topk_accuracies ................ [] |
| retriever_score_scaling ......................... False |
| retriever_seq_length ............................ 256 |
| retro_add_retriever ............................. False |
| retro_attention_gate ............................ 1 |
| retro_cyclic_train_iters ........................ None |
| retro_encoder_attention_dropout ................. 0.1 |
| retro_encoder_hidden_dropout .................... 0.1 |
| retro_encoder_layers ............................ 2 |
| retro_num_neighbors ............................. 2 |
| retro_num_retrieved_chunks ...................... 2 |
| retro_project_dir ............................... None |
| retro_verify_neighbor_count ..................... True |
| rope_scaling_factor ............................. 8.0 |
| rotary_base ..................................... 640000 |
| rotary_interleaved .............................. False |
| rotary_percent .................................. 1.0 |
| rotary_scaling_factor ........................... 1.0 |
| rotary_seq_len_interpolation_factor ............. None |
| run_workload_inspector_server ................... False |
| sample_rate ..................................... 1.0 |
| save ............................................ /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 |
| save_interval ................................... 29 |
| scatter_gather_tensors_in_pipeline .............. True |
| seed ............................................ 1234 |
| seq_length ...................................... 32768 |
| sequence_parallel ............................... True |
| sgd_momentum .................................... 0.9 |
| short_seq_prob .................................. 0.1 |
| skip_data_prepare ............................... False |
| skip_train ...................................... False |
| skipped_train_samples ........................... 0 |
| spec ............................................ ['megatron.core.models.mamba.mamba_layer_specs', 'mamba_moe_stack_spec'] |
| split ........................................... 100,0,0 |
| sqreglu ......................................... False |
| squared_relu .................................... False |
| start_weight_decay .............................. 0.1 |
| straggler_ctrlr_port ............................ 65535 |
| straggler_minmax_count .......................... 1 |
| suggested_communication_unit_size ............... None |
| swiglu .......................................... True |
| swin_backbone_type .............................. tiny |
| te_rng_tracker .................................. False |
| tensor_model_parallel_size ...................... 2 |
| tensorboard_dir ................................. /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/log/2025.09.17-22.24.08_based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 |
| tensorboard_log_interval ........................ 1 |
| tensorboard_queue_size .......................... 1000 |
| test_data_path .................................. None |
| test_mode ....................................... False |
| tiktoken_num_special_tokens ..................... 1000 |
| tiktoken_pattern ................................ None |
| tiktoken_special_tokens ......................... None |
| timing_log_level ................................ 0 |
| timing_log_option ............................... minmax |
| titles_data_path ................................ None |
| tokenizer_model ................................. /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache/models/huggingface/yulan-team/YuLan-Mini |
| tokenizer_type .................................. HuggingFaceTokenizer |
| tp_comm_bootstrap_backend ....................... nccl |
| tp_comm_bulk_dgrad .............................. True |
| tp_comm_bulk_wgrad .............................. True |
| tp_comm_overlap ................................. False |
| tp_comm_overlap_ag .............................. True |
| tp_comm_overlap_cfg ............................. None |
| tp_comm_overlap_rs .............................. True |
| tp_comm_overlap_rs_dgrad ........................ False |
| tp_comm_split_ag ................................ True |
| tp_comm_split_rs ................................ True |
| train_data_path ................................. None |
| train_iters ..................................... None |
| train_samples ................................... 61035 |
| train_sync_interval ............................. None |
| transformer_impl ................................ transformer_engine |
| transformer_pipeline_model_parallel_size ........ 1 |
| untie_embeddings_and_output_weights ............. True |
| use_checkpoint_args ............................. False |
| use_checkpoint_opt_param_scheduler .............. False |
| use_cpu_initialization .......................... None |
| use_custom_fsdp ................................. False |
| use_dist_ckpt ................................... False |
| use_dist_ckpt_deprecated ........................ False |
| use_distributed_optimizer ....................... True |
| use_flash_attn .................................. True |
| use_legacy_models ............................... False |
| use_mp_args_from_checkpoint_args ................ False |
| use_one_sent_docs ............................... False |
| use_persistent_ckpt_worker ...................... False |
| use_precision_aware_optimizer ................... False |
| use_pytorch_profiler ............................ False |
| use_ring_exchange_p2p ........................... False |
| use_rope_scaling ................................ False |
| use_rotary_position_embeddings .................. False |
| use_tokenizer_model_from_checkpoint_args ........ True |
| use_torch_fsdp2 ................................. False |
| use_torch_optimizer_for_cpu_offload ............. False |
| use_tp_pp_dp_mapping ............................ False |
| v_head_dim ...................................... 128 |
| valid_data_path ................................. None |
| variable_seq_lengths ............................ False |
| virtual_pipeline_model_parallel_size ............ None |
| vision_backbone_type ............................ vit |
| vision_pretraining .............................. False |
| vision_pretraining_type ......................... classify |
| vocab_extra_ids ................................. 0 |
| vocab_file ...................................... None |
| vocab_size ...................................... None |
| wandb_exp_name .................................. |
| wandb_project ................................... |
| wandb_save_dir .................................. |
| weight_decay .................................... 0.1 |
| weight_decay_incr_style ......................... constant |
| wgrad_deferral_limit ............................ 0 |
| window_size ..................................... None |
| world_size ...................................... 8 |
| yaml_cfg ........................................ None |
| -------------------- end of arguments --------------------- |
| INFO:megatron.core.num_microbatches_calculator:setting number of microbatches to constant 512 |
| > building HuggingFaceTokenizer tokenizer ... |
| > padded vocab (size: 99000) with 72 dummy tokens (new size: 99072) |
| WARNING:megatron.core.rerun_state_machine:RerunStateMachine initialized in mode RerunMode.DISABLED |
| > initializing torch distributed ... |
| > setting tensorboard ... |
| WARNING: one_logger package is required to enable e2e metrics tracking. please go to https://confluence.nvidia.com/display/MLWFO/Package+Repositories for details to install it |
| > initialized tensor model parallel with size 2 |
| > initialized pipeline model parallel with size 1 |
| > setting random seeds to 1234 ... |
| > compiling dataset index builder ... |
| [rank6]:[W917 22:25:26.615407181 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 6] using GPU 6 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device. |
| [rank4]:[W917 22:25:26.616202999 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 4] using GPU 4 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device. |
| make: Entering directory '/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/datasets' |
| [rank2]:[W917 22:25:26.616462451 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 2] using GPU 2 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device. |
| make: Nothing to be done for 'default'. |
| make: Leaving directory '/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/datasets' |
| >>> done with dataset index builder. Compilation time: 0.241 seconds |
| WARNING: constraints for invoking optimized fused softmax kernel are not met. We default back to unfused kernel invocations. |
| > compiling and loading fused kernels ... |
| [rank0]:[W917 22:25:27.015701688 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 0] using GPU 0 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device. |
| [rank5]:[W917 22:25:27.052206707 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 5] using GPU 5 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device. |
| [rank1]:[W917 22:25:27.054241606 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 1] using GPU 1 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device. |
| [rank3]:[W917 22:25:27.055513460 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 3] using GPU 3 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device. |
| [rank7]:[W917 22:25:27.057617657 ProcessGroupNCCL.cpp:4715] [PG ID 0 PG GUID 0 Rank 7] using GPU 7 as device used by this process is currently unknown. This can potentially cause a hang if this rank to GPU mapping is incorrect. You can pecify device_id in init_process_group() to force use of a particular device. |
| >>> done with compiling and loading fused kernels. Compilation time: 4.598 seconds |
| time to initialize megatron (seconds): 62.726 |
| [after megatron is initialized] datetime: 2025-09-17 22:26:08 |
| building Mamba model ... |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed. |
| warnings.warn( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed. |
| warnings.warn( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed. |
| warnings.warn( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed. |
| warnings.warn( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed. |
| warnings.warn( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed. |
| warnings.warn( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed. |
| warnings.warn( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/transformer/transformer_config.py:788: UserWarning: If you are using transformer_engine as the transformer implementation, the core_attn is from transformer_engine and may be the fused version. For fused attention, you have no need to set 'core_attn' to recompute. Please check that the core_attn recompute is really needed. |
| warnings.warn( |
| INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:Using hybrid override pattern |
| INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:Warning: overriding pattern A with pattern B |
| INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:A: M-M-M-M--M-M-*M-M-M-M-M--M-*M-M-M-M-M-M--*M-M-M-M-M-M-M-*-M-M-M-M-M-M-*M--M-M-M-M-M-*M-M--M-M-M-M-*M-M-M--M-M-M- |
| INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:B: *-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M- |
| INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:Hybrid allocation (M is mamba, * is attention, - is mlp): |
| INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M-*-M-M-M-M-M-M-M- |
| INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:7 attention layers in 112 total layers. |
| INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:Target attention ratio: 0.06. Actual attention ratio: 0.06. |
| INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:56 mlp layers in 112 total layers. |
| INFO:megatron.core.ssm.mamba_hybrid_layer_allocation:Target mlp ratio: 0.50. Actual mlp ratio: 0.50. |
| - decoder.layers.0.self_attention.linear_proj.weight: 1843200 |
| - decoder.layers.0.self_attention.linear_qkv.layer_norm_weight: 1920 |
| - decoder.layers.0.self_attention.linear_qkv.weight: 2580480 |
| - decoder.layers.0.self_attention.linear_qkv.bias: 1344 |
| == params layer 0: 4426944 |
| - decoder.layers.1.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.1.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.1.mlp.linear_fc2.weight: 4608000 |
| == params layer 1: 13825920 |
| - decoder.layers.2.mixer.dt_bias: 15 |
| - decoder.layers.2.mixer.A_log: 15 |
| - decoder.layers.2.mixer.D: 15 |
| - decoder.layers.2.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.2.mixer.in_proj.weight: 7401600 |
| - decoder.layers.2.mixer.conv1d.weight: 11520 |
| - decoder.layers.2.mixer.conv1d.bias: 2880 |
| - decoder.layers.2.mixer.norm.weight: 960 |
| - decoder.layers.2.mixer.out_proj.weight: 1843200 |
| == params layer 2: 9262125 |
| - decoder.layers.3.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.3.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.3.mlp.linear_fc2.weight: 4608000 |
| == params layer 3: 13825920 |
| - decoder.layers.4.mixer.dt_bias: 15 |
| - decoder.layers.4.mixer.A_log: 15 |
| - decoder.layers.4.mixer.D: 15 |
| - decoder.layers.4.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.4.mixer.in_proj.weight: 7401600 |
| - decoder.layers.4.mixer.conv1d.weight: 11520 |
| - decoder.layers.4.mixer.conv1d.bias: 2880 |
| - decoder.layers.4.mixer.norm.weight: 960 |
| - decoder.layers.4.mixer.out_proj.weight: 1843200 |
| == params layer 4: 9262125 |
| - decoder.layers.5.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.5.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.5.mlp.linear_fc2.weight: 4608000 |
| == params layer 5: 13825920 |
| - decoder.layers.6.mixer.dt_bias: 15 |
| - decoder.layers.6.mixer.A_log: 15 |
| - decoder.layers.6.mixer.D: 15 |
| - decoder.layers.6.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.6.mixer.in_proj.weight: 7401600 |
| - decoder.layers.6.mixer.conv1d.weight: 11520 |
| - decoder.layers.6.mixer.conv1d.bias: 2880 |
| - decoder.layers.6.mixer.norm.weight: 960 |
| - decoder.layers.6.mixer.out_proj.weight: 1843200 |
| == params layer 6: 9262125 |
| - decoder.layers.7.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.7.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.7.mlp.linear_fc2.weight: 4608000 |
| == params layer 7: 13825920 |
| - decoder.layers.8.mixer.dt_bias: 15 |
| - decoder.layers.8.mixer.A_log: 15 |
| - decoder.layers.8.mixer.D: 15 |
| - decoder.layers.8.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.8.mixer.in_proj.weight: 7401600 |
| - decoder.layers.8.mixer.conv1d.weight: 11520 |
| - decoder.layers.8.mixer.conv1d.bias: 2880 |
| - decoder.layers.8.mixer.norm.weight: 960 |
| - decoder.layers.8.mixer.out_proj.weight: 1843200 |
| == params layer 8: 9262125 |
| - decoder.layers.9.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.9.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.9.mlp.linear_fc2.weight: 4608000 |
| == params layer 9: 13825920 |
| - decoder.layers.10.mixer.dt_bias: 15 |
| - decoder.layers.10.mixer.A_log: 15 |
| - decoder.layers.10.mixer.D: 15 |
| - decoder.layers.10.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.10.mixer.in_proj.weight: 7401600 |
| - decoder.layers.10.mixer.conv1d.weight: 11520 |
| - decoder.layers.10.mixer.conv1d.bias: 2880 |
| - decoder.layers.10.mixer.norm.weight: 960 |
| - decoder.layers.10.mixer.out_proj.weight: 1843200 |
| == params layer 10: 9262125 |
| - decoder.layers.11.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.11.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.11.mlp.linear_fc2.weight: 4608000 |
| == params layer 11: 13825920 |
| - decoder.layers.12.mixer.dt_bias: 15 |
| - decoder.layers.12.mixer.A_log: 15 |
| - decoder.layers.12.mixer.D: 15 |
| - decoder.layers.12.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.12.mixer.in_proj.weight: 7401600 |
| - decoder.layers.12.mixer.conv1d.weight: 11520 |
| - decoder.layers.12.mixer.conv1d.bias: 2880 |
| - decoder.layers.12.mixer.norm.weight: 960 |
| - decoder.layers.12.mixer.out_proj.weight: 1843200 |
| == params layer 12: 9262125 |
| - decoder.layers.13.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.13.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.13.mlp.linear_fc2.weight: 4608000 |
| == params layer 13: 13825920 |
| - decoder.layers.14.mixer.dt_bias: 15 |
| - decoder.layers.14.mixer.A_log: 15 |
| - decoder.layers.14.mixer.D: 15 |
| - decoder.layers.14.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.14.mixer.in_proj.weight: 7401600 |
| - decoder.layers.14.mixer.conv1d.weight: 11520 |
| - decoder.layers.14.mixer.conv1d.bias: 2880 |
| - decoder.layers.14.mixer.norm.weight: 960 |
| - decoder.layers.14.mixer.out_proj.weight: 1843200 |
| == params layer 14: 9262125 |
| - decoder.layers.15.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.15.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.15.mlp.linear_fc2.weight: 4608000 |
| == params layer 15: 13825920 |
| - decoder.layers.16.self_attention.linear_proj.weight: 1843200 |
| - decoder.layers.16.self_attention.linear_qkv.layer_norm_weight: 1920 |
| - decoder.layers.16.self_attention.linear_qkv.weight: 2580480 |
| - decoder.layers.16.self_attention.linear_qkv.bias: 1344 |
| == params layer 16: 4426944 |
| - decoder.layers.17.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.17.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.17.mlp.linear_fc2.weight: 4608000 |
| == params layer 17: 13825920 |
| - decoder.layers.18.mixer.dt_bias: 15 |
| - decoder.layers.18.mixer.A_log: 15 |
| - decoder.layers.18.mixer.D: 15 |
| - decoder.layers.18.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.18.mixer.in_proj.weight: 7401600 |
| - decoder.layers.18.mixer.conv1d.weight: 11520 |
| - decoder.layers.18.mixer.conv1d.bias: 2880 |
| - decoder.layers.18.mixer.norm.weight: 960 |
| - decoder.layers.18.mixer.out_proj.weight: 1843200 |
| == params layer 18: 9262125 |
| - decoder.layers.19.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.19.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.19.mlp.linear_fc2.weight: 4608000 |
| == params layer 19: 13825920 |
| - decoder.layers.20.mixer.dt_bias: 15 |
| - decoder.layers.20.mixer.A_log: 15 |
| - decoder.layers.20.mixer.D: 15 |
| - decoder.layers.20.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.20.mixer.in_proj.weight: 7401600 |
| - decoder.layers.20.mixer.conv1d.weight: 11520 |
| - decoder.layers.20.mixer.conv1d.bias: 2880 |
| - decoder.layers.20.mixer.norm.weight: 960 |
| - decoder.layers.20.mixer.out_proj.weight: 1843200 |
| == params layer 20: 9262125 |
| - decoder.layers.21.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.21.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.21.mlp.linear_fc2.weight: 4608000 |
| == params layer 21: 13825920 |
| - decoder.layers.22.mixer.dt_bias: 15 |
| - decoder.layers.22.mixer.A_log: 15 |
| - decoder.layers.22.mixer.D: 15 |
| - decoder.layers.22.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.22.mixer.in_proj.weight: 7401600 |
| - decoder.layers.22.mixer.conv1d.weight: 11520 |
| - decoder.layers.22.mixer.conv1d.bias: 2880 |
| - decoder.layers.22.mixer.norm.weight: 960 |
| - decoder.layers.22.mixer.out_proj.weight: 1843200 |
| == params layer 22: 9262125 |
| - decoder.layers.23.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.23.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.23.mlp.linear_fc2.weight: 4608000 |
| == params layer 23: 13825920 |
| - decoder.layers.24.mixer.dt_bias: 15 |
| - decoder.layers.24.mixer.A_log: 15 |
| - decoder.layers.24.mixer.D: 15 |
| - decoder.layers.24.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.24.mixer.in_proj.weight: 7401600 |
| - decoder.layers.24.mixer.conv1d.weight: 11520 |
| - decoder.layers.24.mixer.conv1d.bias: 2880 |
| - decoder.layers.24.mixer.norm.weight: 960 |
| - decoder.layers.24.mixer.out_proj.weight: 1843200 |
| == params layer 24: 9262125 |
| - decoder.layers.25.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.25.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.25.mlp.linear_fc2.weight: 4608000 |
| == params layer 25: 13825920 |
| - decoder.layers.26.mixer.dt_bias: 15 |
| - decoder.layers.26.mixer.A_log: 15 |
| - decoder.layers.26.mixer.D: 15 |
| - decoder.layers.26.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.26.mixer.in_proj.weight: 7401600 |
| - decoder.layers.26.mixer.conv1d.weight: 11520 |
| - decoder.layers.26.mixer.conv1d.bias: 2880 |
| - decoder.layers.26.mixer.norm.weight: 960 |
| - decoder.layers.26.mixer.out_proj.weight: 1843200 |
| == params layer 26: 9262125 |
| - decoder.layers.27.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.27.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.27.mlp.linear_fc2.weight: 4608000 |
| == params layer 27: 13825920 |
| - decoder.layers.28.mixer.dt_bias: 15 |
| - decoder.layers.28.mixer.A_log: 15 |
| - decoder.layers.28.mixer.D: 15 |
| - decoder.layers.28.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.28.mixer.in_proj.weight: 7401600 |
| - decoder.layers.28.mixer.conv1d.weight: 11520 |
| - decoder.layers.28.mixer.conv1d.bias: 2880 |
| - decoder.layers.28.mixer.norm.weight: 960 |
| - decoder.layers.28.mixer.out_proj.weight: 1843200 |
| == params layer 28: 9262125 |
| - decoder.layers.29.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.29.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.29.mlp.linear_fc2.weight: 4608000 |
| == params layer 29: 13825920 |
| - decoder.layers.30.mixer.dt_bias: 15 |
| - decoder.layers.30.mixer.A_log: 15 |
| - decoder.layers.30.mixer.D: 15 |
| - decoder.layers.30.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.30.mixer.in_proj.weight: 7401600 |
| - decoder.layers.30.mixer.conv1d.weight: 11520 |
| - decoder.layers.30.mixer.conv1d.bias: 2880 |
| - decoder.layers.30.mixer.norm.weight: 960 |
| - decoder.layers.30.mixer.out_proj.weight: 1843200 |
| == params layer 30: 9262125 |
| - decoder.layers.31.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.31.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.31.mlp.linear_fc2.weight: 4608000 |
| == params layer 31: 13825920 |
| - decoder.layers.32.self_attention.linear_proj.weight: 1843200 |
| - decoder.layers.32.self_attention.linear_qkv.layer_norm_weight: 1920 |
| - decoder.layers.32.self_attention.linear_qkv.weight: 2580480 |
| - decoder.layers.32.self_attention.linear_qkv.bias: 1344 |
| == params layer 32: 4426944 |
| - decoder.layers.33.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.33.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.33.mlp.linear_fc2.weight: 4608000 |
| == params layer 33: 13825920 |
| - decoder.layers.34.mixer.dt_bias: 15 |
| - decoder.layers.34.mixer.A_log: 15 |
| - decoder.layers.34.mixer.D: 15 |
| - decoder.layers.34.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.34.mixer.in_proj.weight: 7401600 |
| - decoder.layers.34.mixer.conv1d.weight: 11520 |
| - decoder.layers.34.mixer.conv1d.bias: 2880 |
| - decoder.layers.34.mixer.norm.weight: 960 |
| - decoder.layers.34.mixer.out_proj.weight: 1843200 |
| == params layer 34: 9262125 |
| - decoder.layers.35.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.35.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.35.mlp.linear_fc2.weight: 4608000 |
| == params layer 35: 13825920 |
| - decoder.layers.36.mixer.dt_bias: 15 |
| - decoder.layers.36.mixer.A_log: 15 |
| - decoder.layers.36.mixer.D: 15 |
| - decoder.layers.36.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.36.mixer.in_proj.weight: 7401600 |
| - decoder.layers.36.mixer.conv1d.weight: 11520 |
| - decoder.layers.36.mixer.conv1d.bias: 2880 |
| - decoder.layers.36.mixer.norm.weight: 960 |
| - decoder.layers.36.mixer.out_proj.weight: 1843200 |
| == params layer 36: 9262125 |
| - decoder.layers.37.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.37.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.37.mlp.linear_fc2.weight: 4608000 |
| == params layer 37: 13825920 |
| - decoder.layers.38.mixer.dt_bias: 15 |
| - decoder.layers.38.mixer.A_log: 15 |
| - decoder.layers.38.mixer.D: 15 |
| - decoder.layers.38.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.38.mixer.in_proj.weight: 7401600 |
| - decoder.layers.38.mixer.conv1d.weight: 11520 |
| - decoder.layers.38.mixer.conv1d.bias: 2880 |
| - decoder.layers.38.mixer.norm.weight: 960 |
| - decoder.layers.38.mixer.out_proj.weight: 1843200 |
| == params layer 38: 9262125 |
| - decoder.layers.39.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.39.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.39.mlp.linear_fc2.weight: 4608000 |
| == params layer 39: 13825920 |
| - decoder.layers.40.mixer.dt_bias: 15 |
| - decoder.layers.40.mixer.A_log: 15 |
| - decoder.layers.40.mixer.D: 15 |
| - decoder.layers.40.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.40.mixer.in_proj.weight: 7401600 |
| - decoder.layers.40.mixer.conv1d.weight: 11520 |
| - decoder.layers.40.mixer.conv1d.bias: 2880 |
| - decoder.layers.40.mixer.norm.weight: 960 |
| - decoder.layers.40.mixer.out_proj.weight: 1843200 |
| == params layer 40: 9262125 |
| - decoder.layers.41.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.41.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.41.mlp.linear_fc2.weight: 4608000 |
| == params layer 41: 13825920 |
| - decoder.layers.42.mixer.dt_bias: 15 |
| - decoder.layers.42.mixer.A_log: 15 |
| - decoder.layers.42.mixer.D: 15 |
| - decoder.layers.42.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.42.mixer.in_proj.weight: 7401600 |
| - decoder.layers.42.mixer.conv1d.weight: 11520 |
| - decoder.layers.42.mixer.conv1d.bias: 2880 |
| - decoder.layers.42.mixer.norm.weight: 960 |
| - decoder.layers.42.mixer.out_proj.weight: 1843200 |
| == params layer 42: 9262125 |
| - decoder.layers.43.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.43.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.43.mlp.linear_fc2.weight: 4608000 |
| > number of parameters on (tensor, pipeline) model parallel rank (1, 0): 1449304413 |
| == params layer 43: 13825920 |
| - decoder.layers.44.mixer.dt_bias: 15 |
| - decoder.layers.44.mixer.A_log: 15 |
| - decoder.layers.44.mixer.D: 15 |
| - decoder.layers.44.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.44.mixer.in_proj.weight: 7401600 |
| - decoder.layers.44.mixer.conv1d.weight: 11520 |
| - decoder.layers.44.mixer.conv1d.bias: 2880 |
| - decoder.layers.44.mixer.norm.weight: 960 |
| - decoder.layers.44.mixer.out_proj.weight: 1843200 |
| == params layer 44: 9262125 |
| - decoder.layers.45.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.45.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.45.mlp.linear_fc2.weight: 4608000 |
| == params layer 45: 13825920 |
| - decoder.layers.46.mixer.dt_bias: 15 |
| - decoder.layers.46.mixer.A_log: 15 |
| - decoder.layers.46.mixer.D: 15 |
| - decoder.layers.46.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.46.mixer.in_proj.weight: 7401600 |
| - decoder.layers.46.mixer.conv1d.weight: 11520 |
| - decoder.layers.46.mixer.conv1d.bias: 2880 |
| - decoder.layers.46.mixer.norm.weight: 960 |
| - decoder.layers.46.mixer.out_proj.weight: 1843200 |
| == params layer 46: 9262125 |
| - decoder.layers.47.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.47.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.47.mlp.linear_fc2.weight: 4608000 |
| == params layer 47: 13825920 |
| - decoder.layers.48.self_attention.linear_proj.weight: 1843200 |
| - decoder.layers.48.self_attention.linear_qkv.layer_norm_weight: 1920 |
| - decoder.layers.48.self_attention.linear_qkv.weight: 2580480 |
| - decoder.layers.48.self_attention.linear_qkv.bias: 1344 |
| == params layer 48: 4426944 |
| - decoder.layers.49.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.49.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.49.mlp.linear_fc2.weight: 4608000 |
| == params layer 49: 13825920 |
| - decoder.layers.50.mixer.dt_bias: 15 |
| - decoder.layers.50.mixer.A_log: 15 |
| - decoder.layers.50.mixer.D: 15 |
| - decoder.layers.50.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.50.mixer.in_proj.weight: 7401600 |
| - decoder.layers.50.mixer.conv1d.weight: 11520 |
| - decoder.layers.50.mixer.conv1d.bias: 2880 |
| - decoder.layers.50.mixer.norm.weight: 960 |
| - decoder.layers.50.mixer.out_proj.weight: 1843200 |
| == params layer 50: 9262125 |
| - decoder.layers.51.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.51.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.51.mlp.linear_fc2.weight: 4608000 |
| == params layer 51: 13825920 |
| - decoder.layers.52.mixer.dt_bias: 15 |
| - decoder.layers.52.mixer.A_log: 15 |
| - decoder.layers.52.mixer.D: 15 |
| - decoder.layers.52.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.52.mixer.in_proj.weight: 7401600 |
| - decoder.layers.52.mixer.conv1d.weight: 11520 |
| - decoder.layers.52.mixer.conv1d.bias: 2880 |
| - decoder.layers.52.mixer.norm.weight: 960 |
| - decoder.layers.52.mixer.out_proj.weight: 1843200 |
| == params layer 52: 9262125 |
| - decoder.layers.53.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.53.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.53.mlp.linear_fc2.weight: 4608000 |
| == params layer 53: 13825920 |
| - decoder.layers.54.mixer.dt_bias: 15 |
| - decoder.layers.54.mixer.A_log: 15 |
| - decoder.layers.54.mixer.D: 15 |
| - decoder.layers.54.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.54.mixer.in_proj.weight: 7401600 |
| - decoder.layers.54.mixer.conv1d.weight: 11520 |
| - decoder.layers.54.mixer.conv1d.bias: 2880 |
| - decoder.layers.54.mixer.norm.weight: 960 |
| - decoder.layers.54.mixer.out_proj.weight: 1843200 |
| == params layer 54: 9262125 |
| - decoder.layers.55.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.55.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.55.mlp.linear_fc2.weight: 4608000 |
| == params layer 55: 13825920 |
| - decoder.layers.56.mixer.dt_bias: 15 |
| - decoder.layers.56.mixer.A_log: 15 |
| - decoder.layers.56.mixer.D: 15 |
| - decoder.layers.56.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.56.mixer.in_proj.weight: 7401600 |
| - decoder.layers.56.mixer.conv1d.weight: 11520 |
| - decoder.layers.56.mixer.conv1d.bias: 2880 |
| - decoder.layers.56.mixer.norm.weight: 960 |
| - decoder.layers.56.mixer.out_proj.weight: 1843200 |
| == params layer 56: 9262125 |
| - decoder.layers.57.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.57.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.57.mlp.linear_fc2.weight: 4608000 |
| == params layer 57: 13825920 |
| - decoder.layers.58.mixer.dt_bias: 15 |
| - decoder.layers.58.mixer.A_log: 15 |
| - decoder.layers.58.mixer.D: 15 |
| - decoder.layers.58.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.58.mixer.in_proj.weight: 7401600 |
| - decoder.layers.58.mixer.conv1d.weight: 11520 |
| - decoder.layers.58.mixer.conv1d.bias: 2880 |
| - decoder.layers.58.mixer.norm.weight: 960 |
| - decoder.layers.58.mixer.out_proj.weight: 1843200 |
| == params layer 58: 9262125 |
| - decoder.layers.59.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.59.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.59.mlp.linear_fc2.weight: 4608000 |
| == params layer 59: 13825920 |
| - decoder.layers.60.mixer.dt_bias: 15 |
| - decoder.layers.60.mixer.A_log: 15 |
| - decoder.layers.60.mixer.D: 15 |
| - decoder.layers.60.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.60.mixer.in_proj.weight: 7401600 |
| - decoder.layers.60.mixer.conv1d.weight: 11520 |
| - decoder.layers.60.mixer.conv1d.bias: 2880 |
| - decoder.layers.60.mixer.norm.weight: 960 |
| - decoder.layers.60.mixer.out_proj.weight: 1843200 |
| == params layer 60: 9262125 |
| - decoder.layers.61.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.61.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.61.mlp.linear_fc2.weight: 4608000 |
| == params layer 61: 13825920 |
| - decoder.layers.62.mixer.dt_bias: 15 |
| - decoder.layers.62.mixer.A_log: 15 |
| - decoder.layers.62.mixer.D: 15 |
| - decoder.layers.62.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.62.mixer.in_proj.weight: 7401600 |
| - decoder.layers.62.mixer.conv1d.weight: 11520 |
| - decoder.layers.62.mixer.conv1d.bias: 2880 |
| - decoder.layers.62.mixer.norm.weight: 960 |
| - decoder.layers.62.mixer.out_proj.weight: 1843200 |
| == params layer 62: 9262125 |
| - decoder.layers.63.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.63.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.63.mlp.linear_fc2.weight: 4608000 |
| == params layer 63: 13825920 |
| - decoder.layers.64.self_attention.linear_proj.weight: 1843200 |
| - decoder.layers.64.self_attention.linear_qkv.layer_norm_weight: 1920 |
| - decoder.layers.64.self_attention.linear_qkv.weight: 2580480 |
| - decoder.layers.64.self_attention.linear_qkv.bias: 1344 |
| == params layer 64: 4426944 |
| - decoder.layers.65.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.65.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.65.mlp.linear_fc2.weight: 4608000 |
| == params layer 65: 13825920 |
| - decoder.layers.66.mixer.dt_bias: 15 |
| - decoder.layers.66.mixer.A_log: 15 |
| - decoder.layers.66.mixer.D: 15 |
| - decoder.layers.66.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.66.mixer.in_proj.weight: 7401600 |
| - decoder.layers.66.mixer.conv1d.weight: 11520 |
| - decoder.layers.66.mixer.conv1d.bias: 2880 |
| - decoder.layers.66.mixer.norm.weight: 960 |
| - decoder.layers.66.mixer.out_proj.weight: 1843200 |
| == params layer 66: 9262125 |
| - decoder.layers.67.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.67.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.67.mlp.linear_fc2.weight: 4608000 |
| == params layer 67: 13825920 |
| - decoder.layers.68.mixer.dt_bias: 15 |
| - decoder.layers.68.mixer.A_log: 15 |
| - decoder.layers.68.mixer.D: 15 |
| - decoder.layers.68.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.68.mixer.in_proj.weight: 7401600 |
| - decoder.layers.68.mixer.conv1d.weight: 11520 |
| - decoder.layers.68.mixer.conv1d.bias: 2880 |
| - decoder.layers.68.mixer.norm.weight: 960 |
| - decoder.layers.68.mixer.out_proj.weight: 1843200 |
| == params layer 68: 9262125 |
| - decoder.layers.69.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.69.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.69.mlp.linear_fc2.weight: 4608000 |
| == params layer 69: 13825920 |
| - decoder.layers.70.mixer.dt_bias: 15 |
| - decoder.layers.70.mixer.A_log: 15 |
| - decoder.layers.70.mixer.D: 15 |
| - decoder.layers.70.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.70.mixer.in_proj.weight: 7401600 |
| - decoder.layers.70.mixer.conv1d.weight: 11520 |
| - decoder.layers.70.mixer.conv1d.bias: 2880 |
| - decoder.layers.70.mixer.norm.weight: 960 |
| - decoder.layers.70.mixer.out_proj.weight: 1843200 |
| == params layer 70: 9262125 |
| - decoder.layers.71.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.71.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.71.mlp.linear_fc2.weight: 4608000 |
| == params layer 71: 13825920 |
| - decoder.layers.72.mixer.dt_bias: 15 |
| - decoder.layers.72.mixer.A_log: 15 |
| - decoder.layers.72.mixer.D: 15 |
| - decoder.layers.72.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.72.mixer.in_proj.weight: 7401600 |
| - decoder.layers.72.mixer.conv1d.weight: 11520 |
| - decoder.layers.72.mixer.conv1d.bias: 2880 |
| - decoder.layers.72.mixer.norm.weight: 960 |
| - decoder.layers.72.mixer.out_proj.weight: 1843200 |
| == params layer 72: 9262125 |
| - decoder.layers.73.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.73.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.73.mlp.linear_fc2.weight: 4608000 |
| == params layer 73: 13825920 |
| - decoder.layers.74.mixer.dt_bias: 15 |
| - decoder.layers.74.mixer.A_log: 15 |
| - decoder.layers.74.mixer.D: 15 |
| - decoder.layers.74.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.74.mixer.in_proj.weight: 7401600 |
| - decoder.layers.74.mixer.conv1d.weight: 11520 |
| - decoder.layers.74.mixer.conv1d.bias: 2880 |
| - decoder.layers.74.mixer.norm.weight: 960 |
| - decoder.layers.74.mixer.out_proj.weight: 1843200 |
| == params layer 74: 9262125 |
| - decoder.layers.75.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.75.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.75.mlp.linear_fc2.weight: 4608000 |
| == params layer 75: 13825920 |
| - decoder.layers.76.mixer.dt_bias: 15 |
| - decoder.layers.76.mixer.A_log: 15 |
| - decoder.layers.76.mixer.D: 15 |
| - decoder.layers.76.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.76.mixer.in_proj.weight: 7401600 |
| - decoder.layers.76.mixer.conv1d.weight: 11520 |
| - decoder.layers.76.mixer.conv1d.bias: 2880 |
| - decoder.layers.76.mixer.norm.weight: 960 |
| - decoder.layers.76.mixer.out_proj.weight: 1843200 |
| == params layer 76: 9262125 |
| - decoder.layers.77.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.77.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.77.mlp.linear_fc2.weight: 4608000 |
| == params layer 77: 13825920 |
| - decoder.layers.78.mixer.dt_bias: 15 |
| - decoder.layers.78.mixer.A_log: 15 |
| - decoder.layers.78.mixer.D: 15 |
| - decoder.layers.78.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.78.mixer.in_proj.weight: 7401600 |
| - decoder.layers.78.mixer.conv1d.weight: 11520 |
| - decoder.layers.78.mixer.conv1d.bias: 2880 |
| - decoder.layers.78.mixer.norm.weight: 960 |
| - decoder.layers.78.mixer.out_proj.weight: 1843200 |
| == params layer 78: 9262125 |
| - decoder.layers.79.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.79.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.79.mlp.linear_fc2.weight: 4608000 |
| == params layer 79: 13825920 |
| - decoder.layers.80.self_attention.linear_proj.weight: 1843200 |
| - decoder.layers.80.self_attention.linear_qkv.layer_norm_weight: 1920 |
| - decoder.layers.80.self_attention.linear_qkv.weight: 2580480 |
| - decoder.layers.80.self_attention.linear_qkv.bias: 1344 |
| == params layer 80: 4426944 |
| - decoder.layers.81.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.81.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.81.mlp.linear_fc2.weight: 4608000 |
| == params layer 81: 13825920 |
| - decoder.layers.82.mixer.dt_bias: 15 |
| - decoder.layers.82.mixer.A_log: 15 |
| - decoder.layers.82.mixer.D: 15 |
| - decoder.layers.82.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.82.mixer.in_proj.weight: 7401600 |
| - decoder.layers.82.mixer.conv1d.weight: 11520 |
| - decoder.layers.82.mixer.conv1d.bias: 2880 |
| - decoder.layers.82.mixer.norm.weight: 960 |
| - decoder.layers.82.mixer.out_proj.weight: 1843200 |
| == params layer 82: 9262125 |
| - decoder.layers.83.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.83.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.83.mlp.linear_fc2.weight: 4608000 |
| == params layer 83: 13825920 |
| - decoder.layers.84.mixer.dt_bias: 15 |
| - decoder.layers.84.mixer.A_log: 15 |
| - decoder.layers.84.mixer.D: 15 |
| - decoder.layers.84.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.84.mixer.in_proj.weight: 7401600 |
| - decoder.layers.84.mixer.conv1d.weight: 11520 |
| - decoder.layers.84.mixer.conv1d.bias: 2880 |
| - decoder.layers.84.mixer.norm.weight: 960 |
| - decoder.layers.84.mixer.out_proj.weight: 1843200 |
| == params layer 84: 9262125 |
| - decoder.layers.85.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.85.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.85.mlp.linear_fc2.weight: 4608000 |
| == params layer 85: 13825920 |
| - decoder.layers.86.mixer.dt_bias: 15 |
| - decoder.layers.86.mixer.A_log: 15 |
| - decoder.layers.86.mixer.D: 15 |
| - decoder.layers.86.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.86.mixer.in_proj.weight: 7401600 |
| - decoder.layers.86.mixer.conv1d.weight: 11520 |
| - decoder.layers.86.mixer.conv1d.bias: 2880 |
| - decoder.layers.86.mixer.norm.weight: 960 |
| - decoder.layers.86.mixer.out_proj.weight: 1843200 |
| == params layer 86: 9262125 |
| - decoder.layers.87.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.87.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.87.mlp.linear_fc2.weight: 4608000 |
| == params layer 87: 13825920 |
| - decoder.layers.88.mixer.dt_bias: 15 |
| - decoder.layers.88.mixer.A_log: 15 |
| - decoder.layers.88.mixer.D: 15 |
| - decoder.layers.88.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.88.mixer.in_proj.weight: 7401600 |
| - decoder.layers.88.mixer.conv1d.weight: 11520 |
| - decoder.layers.88.mixer.conv1d.bias: 2880 |
| - decoder.layers.88.mixer.norm.weight: 960 |
| - decoder.layers.88.mixer.out_proj.weight: 1843200 |
| == params layer 88: 9262125 |
| - decoder.layers.89.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.89.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.89.mlp.linear_fc2.weight: 4608000 |
| == params layer 89: 13825920 |
| - decoder.layers.90.mixer.dt_bias: 15 |
| - decoder.layers.90.mixer.A_log: 15 |
| - decoder.layers.90.mixer.D: 15 |
| - decoder.layers.90.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.90.mixer.in_proj.weight: 7401600 |
| - decoder.layers.90.mixer.conv1d.weight: 11520 |
| - decoder.layers.90.mixer.conv1d.bias: 2880 |
| - decoder.layers.90.mixer.norm.weight: 960 |
| - decoder.layers.90.mixer.out_proj.weight: 1843200 |
| == params layer 90: 9262125 |
| - decoder.layers.91.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.91.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.91.mlp.linear_fc2.weight: 4608000 |
| == params layer 91: 13825920 |
| - decoder.layers.92.mixer.dt_bias: 15 |
| - decoder.layers.92.mixer.A_log: 15 |
| - decoder.layers.92.mixer.D: 15 |
| - decoder.layers.92.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.92.mixer.in_proj.weight: 7401600 |
| - decoder.layers.92.mixer.conv1d.weight: 11520 |
| - decoder.layers.92.mixer.conv1d.bias: 2880 |
| - decoder.layers.92.mixer.norm.weight: 960 |
| - decoder.layers.92.mixer.out_proj.weight: 1843200 |
| == params layer 92: 9262125 |
| - decoder.layers.93.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.93.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.93.mlp.linear_fc2.weight: 4608000 |
| == params layer 93: 13825920 |
| - decoder.layers.94.mixer.dt_bias: 15 |
| - decoder.layers.94.mixer.A_log: 15 |
| - decoder.layers.94.mixer.D: 15 |
| - decoder.layers.94.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.94.mixer.in_proj.weight: 7401600 |
| - decoder.layers.94.mixer.conv1d.weight: 11520 |
| - decoder.layers.94.mixer.conv1d.bias: 2880 |
| - decoder.layers.94.mixer.norm.weight: 960 |
| - decoder.layers.94.mixer.out_proj.weight: 1843200 |
| == params layer 94: 9262125 |
| - decoder.layers.95.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.95.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.95.mlp.linear_fc2.weight: 4608000 |
| == params layer 95: 13825920 |
| - decoder.layers.96.self_attention.linear_proj.weight: 1843200 |
| - decoder.layers.96.self_attention.linear_qkv.layer_norm_weight: 1920 |
| - decoder.layers.96.self_attention.linear_qkv.weight: 2580480 |
| - decoder.layers.96.self_attention.linear_qkv.bias: 1344 |
| == params layer 96: 4426944 |
| - decoder.layers.97.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.97.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.97.mlp.linear_fc2.weight: 4608000 |
| == params layer 97: 13825920 |
| - decoder.layers.98.mixer.dt_bias: 15 |
| - decoder.layers.98.mixer.A_log: 15 |
| - decoder.layers.98.mixer.D: 15 |
| - decoder.layers.98.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.98.mixer.in_proj.weight: 7401600 |
| - decoder.layers.98.mixer.conv1d.weight: 11520 |
| - decoder.layers.98.mixer.conv1d.bias: 2880 |
| - decoder.layers.98.mixer.norm.weight: 960 |
| - decoder.layers.98.mixer.out_proj.weight: 1843200 |
| == params layer 98: 9262125 |
| - decoder.layers.99.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.99.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.99.mlp.linear_fc2.weight: 4608000 |
| == params layer 99: 13825920 |
| - decoder.layers.100.mixer.dt_bias: 15 |
| - decoder.layers.100.mixer.A_log: 15 |
| - decoder.layers.100.mixer.D: 15 |
| - decoder.layers.100.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.100.mixer.in_proj.weight: 7401600 |
| - decoder.layers.100.mixer.conv1d.weight: 11520 |
| - decoder.layers.100.mixer.conv1d.bias: 2880 |
| - decoder.layers.100.mixer.norm.weight: 960 |
| - decoder.layers.100.mixer.out_proj.weight: 1843200 |
| == params layer 100: 9262125 |
| - decoder.layers.101.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.101.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.101.mlp.linear_fc2.weight: 4608000 |
| == params layer 101: 13825920 |
| - decoder.layers.102.mixer.dt_bias: 15 |
| - decoder.layers.102.mixer.A_log: 15 |
| - decoder.layers.102.mixer.D: 15 |
| - decoder.layers.102.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.102.mixer.in_proj.weight: 7401600 |
| - decoder.layers.102.mixer.conv1d.weight: 11520 |
| - decoder.layers.102.mixer.conv1d.bias: 2880 |
| - decoder.layers.102.mixer.norm.weight: 960 |
| - decoder.layers.102.mixer.out_proj.weight: 1843200 |
| == params layer 102: 9262125 |
| - decoder.layers.103.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.103.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.103.mlp.linear_fc2.weight: 4608000 |
| == params layer 103: 13825920 |
| - decoder.layers.104.mixer.dt_bias: 15 |
| - decoder.layers.104.mixer.A_log: 15 |
| - decoder.layers.104.mixer.D: 15 |
| - decoder.layers.104.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.104.mixer.in_proj.weight: 7401600 |
| - decoder.layers.104.mixer.conv1d.weight: 11520 |
| - decoder.layers.104.mixer.conv1d.bias: 2880 |
| - decoder.layers.104.mixer.norm.weight: 960 |
| - decoder.layers.104.mixer.out_proj.weight: 1843200 |
| == params layer 104: 9262125 |
| - decoder.layers.105.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.105.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.105.mlp.linear_fc2.weight: 4608000 |
| == params layer 105: 13825920 |
| - decoder.layers.106.mixer.dt_bias: 15 |
| - decoder.layers.106.mixer.A_log: 15 |
| - decoder.layers.106.mixer.D: 15 |
| - decoder.layers.106.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.106.mixer.in_proj.weight: 7401600 |
| - decoder.layers.106.mixer.conv1d.weight: 11520 |
| - decoder.layers.106.mixer.conv1d.bias: 2880 |
| - decoder.layers.106.mixer.norm.weight: 960 |
| - decoder.layers.106.mixer.out_proj.weight: 1843200 |
| == params layer 106: 9262125 |
| - decoder.layers.107.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.107.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.107.mlp.linear_fc2.weight: 4608000 |
| == params layer 107: 13825920 |
| - decoder.layers.108.mixer.dt_bias: 15 |
| - decoder.layers.108.mixer.A_log: 15 |
| - decoder.layers.108.mixer.D: 15 |
| - decoder.layers.108.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.108.mixer.in_proj.weight: 7401600 |
| - decoder.layers.108.mixer.conv1d.weight: 11520 |
| - decoder.layers.108.mixer.conv1d.bias: 2880 |
| - decoder.layers.108.mixer.norm.weight: 960 |
| - decoder.layers.108.mixer.out_proj.weight: 1843200 |
| == params layer 108: 9262125 |
| - decoder.layers.109.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.109.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.109.mlp.linear_fc2.weight: 4608000 |
| == params layer 109: 13825920 |
| - decoder.layers.110.mixer.dt_bias: 15 |
| - decoder.layers.110.mixer.A_log: 15 |
| - decoder.layers.110.mixer.D: 15 |
| - decoder.layers.110.mixer.in_proj.layer_norm_weight: 1920 |
| - decoder.layers.110.mixer.in_proj.weight: 7401600 |
| - decoder.layers.110.mixer.conv1d.weight: 11520 |
| - decoder.layers.110.mixer.conv1d.bias: 2880 |
| - decoder.layers.110.mixer.norm.weight: 960 |
| - decoder.layers.110.mixer.out_proj.weight: 1843200 |
| == params layer 110: 9262125 |
| - decoder.layers.111.mlp.linear_fc1.layer_norm_weight: 1920 |
| - decoder.layers.111.mlp.linear_fc1.weight: 9216000 |
| - decoder.layers.111.mlp.linear_fc2.weight: 4608000 |
| == params layer 111: 13825920 |
| > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 1449304413 |
| [DEBUG] freeze_non_mamba: False |
| [DEBUG] freeze_non_mamba: False |
| > number of parameters on (tensor, pipeline) model parallel rank (1, 0): 1449304413 |
| INFO:megatron.core.distributed.distributed_data_parallel:Setting up DistributedDataParallel with config DistributedDataParallelConfig(grad_reduce_in_fp32=True, overlap_grad_reduce=True, overlap_param_gather=True, align_param_gather=False, use_distributed_optimizer=True, num_distributed_optimizer_instances=1, check_for_nan_in_grad=True, check_for_large_grads=False, bucket_size=40000000, pad_buckets_for_high_nccl_busbw=False, average_in_collective=False, fp8_param_gather=False, use_custom_fsdp=False, data_parallel_sharding_strategy='no_shard', gradient_reduce_div_fusion=True, suggested_communication_unit_size=None, preserve_fp32_weights=True, keep_fp8_transpose_cache_when_using_custom_fsdp=False) |
| > number of parameters on (tensor, pipeline) model parallel rank (0, 0): 1449304413 |
| [DEBUG] freeze_non_mamba: False |
| INFO:megatron.core.distributed.param_and_grad_buffer:Number of buckets for gradient all-reduce / reduce-scatter: 30 |
| Params for bucket 1 (95109120 elements, 95109120 padded size): |
| module.output_layer.weight |
| Params for bucket 2 (46176045 elements, 46176256 padded size): |
| module.decoder.layers.109.mlp.linear_fc1.weight |
| module.decoder.layers.108.mixer.out_proj.weight |
| module.decoder.layers.111.mlp.linear_fc1.weight |
| module.decoder.layers.110.mixer.out_proj.weight |
| module.decoder.layers.110.mixer.conv1d.weight |
| module.decoder.layers.110.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.108.mixer.conv1d.weight |
| module.decoder.layers.110.mixer.dt_bias |
| module.decoder.layers.108.mixer.conv1d.bias |
| module.decoder.layers.110.mixer.conv1d.bias |
| module.decoder.layers.109.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.110.mixer.D |
| module.decoder.layers.110.mixer.A_log |
| module.decoder.layers.111.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.108.mixer.norm.weight |
| module.decoder.final_norm.weight |
| module.decoder.layers.111.mlp.linear_fc2.weight |
| module.decoder.layers.110.mixer.norm.weight |
| module.decoder.layers.110.mixer.in_proj.weight |
| module.decoder.layers.109.mlp.linear_fc2.weight |
| module.decoder.layers.108.mixer.in_proj.weight |
| Params for bucket 3 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.106.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.106.mixer.D |
| module.decoder.layers.105.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.108.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.107.mlp.linear_fc1.weight |
| module.decoder.layers.106.mixer.out_proj.weight |
| module.decoder.layers.106.mixer.conv1d.bias |
| module.decoder.layers.106.mixer.dt_bias |
| module.decoder.layers.105.mlp.linear_fc2.weight |
| module.decoder.layers.108.mixer.dt_bias |
| module.decoder.layers.104.mixer.in_proj.weight |
| module.decoder.layers.106.mixer.A_log |
| module.decoder.layers.108.mixer.A_log |
| module.decoder.layers.108.mixer.D |
| module.decoder.layers.107.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.106.mixer.norm.weight |
| module.decoder.layers.105.mlp.linear_fc1.weight |
| module.decoder.layers.104.mixer.out_proj.weight |
| module.decoder.layers.104.mixer.norm.weight |
| module.decoder.layers.104.mixer.conv1d.weight |
| module.decoder.layers.106.mixer.in_proj.weight |
| module.decoder.layers.107.mlp.linear_fc2.weight |
| module.decoder.layers.106.mixer.conv1d.weight |
| module.decoder.layers.104.mixer.conv1d.bias |
| Params for bucket 4 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.102.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.102.mixer.D |
| module.decoder.layers.102.mixer.A_log |
| module.decoder.layers.101.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.104.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.104.mixer.A_log |
| module.decoder.layers.104.mixer.D |
| module.decoder.layers.103.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.102.mixer.in_proj.weight |
| module.decoder.layers.101.mlp.linear_fc2.weight |
| module.decoder.layers.103.mlp.linear_fc2.weight |
| module.decoder.layers.100.mixer.in_proj.weight |
| module.decoder.layers.103.mlp.linear_fc1.weight |
| module.decoder.layers.102.mixer.out_proj.weight |
| module.decoder.layers.102.mixer.norm.weight |
| module.decoder.layers.102.mixer.conv1d.weight |
| module.decoder.layers.101.mlp.linear_fc1.weight |
| module.decoder.layers.100.mixer.out_proj.weight |
| module.decoder.layers.100.mixer.norm.weight |
| module.decoder.layers.100.mixer.conv1d.weight |
| module.decoder.layers.102.mixer.dt_bias |
| module.decoder.layers.104.mixer.dt_bias |
| module.decoder.layers.102.mixer.conv1d.bias |
| module.decoder.layers.100.mixer.conv1d.bias |
| Params for bucket 5 (41342874 elements, 41343232 padded size): |
| module.decoder.layers.98.mixer.A_log |
| module.decoder.layers.100.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.100.mixer.A_log |
| module.decoder.layers.100.mixer.D |
| module.decoder.layers.99.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.98.mixer.norm.weight |
| module.decoder.layers.98.mixer.D |
| module.decoder.layers.96.self_attention.linear_qkv.weight |
| module.decoder.layers.99.mlp.linear_fc2.weight |
| module.decoder.layers.98.mixer.conv1d.bias |
| module.decoder.layers.98.mixer.in_proj.weight |
| module.decoder.layers.97.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.96.self_attention.linear_qkv.bias |
| module.decoder.layers.96.self_attention.linear_proj.weight |
| module.decoder.layers.97.mlp.linear_fc1.weight |
| module.decoder.layers.98.mixer.conv1d.weight |
| module.decoder.layers.99.mlp.linear_fc1.weight |
| module.decoder.layers.98.mixer.out_proj.weight |
| module.decoder.layers.98.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.98.mixer.dt_bias |
| module.decoder.layers.97.mlp.linear_fc2.weight |
| module.decoder.layers.100.mixer.dt_bias |
| module.decoder.layers.96.self_attention.linear_qkv.layer_norm_weight |
| Params for bucket 6 (46174125 elements, 46174336 padded size): |
| module.decoder.layers.95.mlp.linear_fc2.weight |
| module.decoder.layers.94.mixer.norm.weight |
| module.decoder.layers.94.mixer.in_proj.weight |
| module.decoder.layers.92.mixer.norm.weight |
| module.decoder.layers.93.mlp.linear_fc2.weight |
| module.decoder.layers.92.mixer.in_proj.weight |
| module.decoder.layers.93.mlp.linear_fc1.weight |
| module.decoder.layers.94.mixer.out_proj.weight |
| module.decoder.layers.95.mlp.linear_fc1.weight |
| module.decoder.layers.94.mixer.conv1d.weight |
| module.decoder.layers.94.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.92.mixer.conv1d.weight |
| module.decoder.layers.92.mixer.out_proj.weight |
| module.decoder.layers.92.mixer.conv1d.bias |
| module.decoder.layers.94.mixer.dt_bias |
| module.decoder.layers.94.mixer.conv1d.bias |
| module.decoder.layers.93.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.95.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.94.mixer.D |
| module.decoder.layers.94.mixer.A_log |
| Params for bucket 7 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.90.mixer.in_proj.weight |
| module.decoder.layers.91.mlp.linear_fc2.weight |
| module.decoder.layers.90.mixer.norm.weight |
| module.decoder.layers.89.mlp.linear_fc2.weight |
| module.decoder.layers.88.mixer.norm.weight |
| module.decoder.layers.88.mixer.in_proj.weight |
| module.decoder.layers.90.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.90.mixer.conv1d.weight |
| module.decoder.layers.92.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.91.mlp.linear_fc1.weight |
| module.decoder.layers.90.mixer.out_proj.weight |
| module.decoder.layers.89.mlp.linear_fc1.weight |
| module.decoder.layers.88.mixer.out_proj.weight |
| module.decoder.layers.88.mixer.conv1d.weight |
| module.decoder.layers.90.mixer.dt_bias |
| module.decoder.layers.92.mixer.dt_bias |
| module.decoder.layers.90.mixer.conv1d.bias |
| module.decoder.layers.88.mixer.conv1d.bias |
| module.decoder.layers.89.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.90.mixer.A_log |
| module.decoder.layers.92.mixer.D |
| module.decoder.layers.92.mixer.A_log |
| module.decoder.layers.91.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.90.mixer.D |
| Params for bucket 8 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.86.mixer.in_proj.weight |
| module.decoder.layers.87.mlp.linear_fc2.weight |
| module.decoder.layers.86.mixer.norm.weight |
| module.decoder.layers.85.mlp.linear_fc2.weight |
| module.decoder.layers.84.mixer.conv1d.bias |
| module.decoder.layers.84.mixer.conv1d.weight |
| module.decoder.layers.86.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.86.mixer.conv1d.weight |
| module.decoder.layers.85.mlp.linear_fc1.weight |
| module.decoder.layers.88.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.87.mlp.linear_fc1.weight |
| module.decoder.layers.86.mixer.out_proj.weight |
| module.decoder.layers.85.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.84.mixer.out_proj.weight |
| module.decoder.layers.84.mixer.in_proj.weight |
| module.decoder.layers.86.mixer.dt_bias |
| module.decoder.layers.88.mixer.dt_bias |
| module.decoder.layers.86.mixer.conv1d.bias |
| module.decoder.layers.86.mixer.A_log |
| module.decoder.layers.88.mixer.A_log |
| module.decoder.layers.88.mixer.D |
| module.decoder.layers.87.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.86.mixer.D |
| module.decoder.layers.84.mixer.norm.weight |
| Params for bucket 9 (41342874 elements, 41343232 padded size): |
| module.decoder.layers.82.mixer.dt_bias |
| module.decoder.layers.81.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.81.mlp.linear_fc2.weight |
| module.decoder.layers.80.self_attention.linear_qkv.layer_norm_weight |
| module.decoder.layers.82.mixer.D |
| module.decoder.layers.84.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.84.mixer.D |
| module.decoder.layers.83.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.82.mixer.norm.weight |
| module.decoder.layers.82.mixer.A_log |
| module.decoder.layers.80.self_attention.linear_qkv.weight |
| module.decoder.layers.82.mixer.in_proj.weight |
| module.decoder.layers.83.mlp.linear_fc2.weight |
| module.decoder.layers.84.mixer.dt_bias |
| module.decoder.layers.82.mixer.conv1d.bias |
| module.decoder.layers.80.self_attention.linear_qkv.bias |
| module.decoder.layers.80.self_attention.linear_proj.weight |
| module.decoder.layers.82.mixer.conv1d.weight |
| module.decoder.layers.82.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.81.mlp.linear_fc1.weight |
| module.decoder.layers.84.mixer.A_log |
| module.decoder.layers.83.mlp.linear_fc1.weight |
| module.decoder.layers.82.mixer.out_proj.weight |
| Params for bucket 10 (46174125 elements, 46174336 padded size): |
| module.decoder.layers.78.mixer.A_log |
| module.decoder.layers.77.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.79.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.78.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.78.mixer.D |
| module.decoder.layers.77.mlp.linear_fc2.weight |
| module.decoder.layers.79.mlp.linear_fc2.weight |
| module.decoder.layers.78.mixer.in_proj.weight |
| module.decoder.layers.76.mixer.in_proj.weight |
| module.decoder.layers.76.mixer.conv1d.weight |
| module.decoder.layers.76.mixer.out_proj.weight |
| module.decoder.layers.79.mlp.linear_fc1.weight |
| module.decoder.layers.78.mixer.norm.weight |
| module.decoder.layers.78.mixer.out_proj.weight |
| module.decoder.layers.78.mixer.conv1d.weight |
| module.decoder.layers.77.mlp.linear_fc1.weight |
| module.decoder.layers.76.mixer.norm.weight |
| module.decoder.layers.76.mixer.conv1d.bias |
| module.decoder.layers.78.mixer.conv1d.bias |
| module.decoder.layers.78.mixer.dt_bias |
| Params for bucket 11 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.74.mixer.D |
| module.decoder.layers.74.mixer.A_log |
| module.decoder.layers.73.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.76.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.76.mixer.D |
| module.decoder.layers.76.mixer.A_log |
| module.decoder.layers.75.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.74.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.73.mlp.linear_fc2.weight |
| module.decoder.layers.75.mlp.linear_fc2.weight |
| module.decoder.layers.74.mixer.in_proj.weight |
| module.decoder.layers.72.mixer.in_proj.weight |
| module.decoder.layers.73.mlp.linear_fc1.weight |
| module.decoder.layers.75.mlp.linear_fc1.weight |
| module.decoder.layers.74.mixer.out_proj.weight |
| module.decoder.layers.74.mixer.norm.weight |
| module.decoder.layers.74.mixer.conv1d.weight |
| module.decoder.layers.72.mixer.out_proj.weight |
| module.decoder.layers.72.mixer.norm.weight |
| module.decoder.layers.72.mixer.conv1d.weight |
| module.decoder.layers.76.mixer.dt_bias |
| module.decoder.layers.74.mixer.conv1d.bias |
| module.decoder.layers.74.mixer.dt_bias |
| module.decoder.layers.72.mixer.conv1d.bias |
| Params for bucket 12 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.70.mixer.D |
| module.decoder.layers.70.mixer.A_log |
| module.decoder.layers.69.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.72.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.72.mixer.D |
| module.decoder.layers.72.mixer.A_log |
| module.decoder.layers.71.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.70.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.69.mlp.linear_fc2.weight |
| module.decoder.layers.71.mlp.linear_fc2.weight |
| module.decoder.layers.70.mixer.in_proj.weight |
| module.decoder.layers.68.mixer.in_proj.weight |
| module.decoder.layers.69.mlp.linear_fc1.weight |
| module.decoder.layers.71.mlp.linear_fc1.weight |
| module.decoder.layers.70.mixer.out_proj.weight |
| module.decoder.layers.70.mixer.norm.weight |
| module.decoder.layers.70.mixer.conv1d.weight |
| module.decoder.layers.68.mixer.out_proj.weight |
| module.decoder.layers.68.mixer.norm.weight |
| module.decoder.layers.68.mixer.conv1d.weight |
| module.decoder.layers.72.mixer.dt_bias |
| module.decoder.layers.70.mixer.conv1d.bias |
| module.decoder.layers.70.mixer.dt_bias |
| module.decoder.layers.68.mixer.conv1d.bias |
| Params for bucket 13 (41342874 elements, 41343232 padded size): |
| module.decoder.layers.66.mixer.A_log |
| module.decoder.layers.68.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.68.mixer.D |
| module.decoder.layers.68.mixer.A_log |
| module.decoder.layers.67.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.66.mixer.norm.weight |
| module.decoder.layers.66.mixer.D |
| module.decoder.layers.64.self_attention.linear_qkv.weight |
| module.decoder.layers.65.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.67.mlp.linear_fc2.weight |
| module.decoder.layers.66.mixer.conv1d.bias |
| module.decoder.layers.66.mixer.in_proj.weight |
| module.decoder.layers.64.self_attention.linear_qkv.bias |
| module.decoder.layers.64.self_attention.linear_proj.weight |
| module.decoder.layers.65.mlp.linear_fc1.weight |
| module.decoder.layers.67.mlp.linear_fc1.weight |
| module.decoder.layers.66.mixer.out_proj.weight |
| module.decoder.layers.66.mixer.conv1d.weight |
| module.decoder.layers.66.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.65.mlp.linear_fc2.weight |
| module.decoder.layers.66.mixer.dt_bias |
| module.decoder.layers.68.mixer.dt_bias |
| module.decoder.layers.64.self_attention.linear_qkv.layer_norm_weight |
| Params for bucket 14 (46174125 elements, 46174336 padded size): |
| module.decoder.layers.61.mlp.linear_fc1.weight |
| module.decoder.layers.63.mlp.linear_fc2.weight |
| module.decoder.layers.62.mixer.in_proj.weight |
| module.decoder.layers.62.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.60.mixer.conv1d.bias |
| module.decoder.layers.61.mlp.linear_fc2.weight |
| module.decoder.layers.63.mlp.linear_fc1.weight |
| module.decoder.layers.62.mixer.out_proj.weight |
| module.decoder.layers.62.mixer.conv1d.bias |
| module.decoder.layers.62.mixer.conv1d.weight |
| module.decoder.layers.62.mixer.D |
| module.decoder.layers.61.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.62.mixer.dt_bias |
| module.decoder.layers.60.mixer.in_proj.weight |
| module.decoder.layers.60.mixer.out_proj.weight |
| module.decoder.layers.60.mixer.norm.weight |
| module.decoder.layers.62.mixer.norm.weight |
| module.decoder.layers.63.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.62.mixer.A_log |
| module.decoder.layers.60.mixer.conv1d.weight |
| Params for bucket 15 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.58.mixer.dt_bias |
| module.decoder.layers.60.mixer.dt_bias |
| module.decoder.layers.58.mixer.conv1d.bias |
| module.decoder.layers.56.mixer.conv1d.bias |
| module.decoder.layers.58.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.58.mixer.D |
| module.decoder.layers.60.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.60.mixer.A_log |
| module.decoder.layers.60.mixer.D |
| module.decoder.layers.59.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.58.mixer.A_log |
| module.decoder.layers.57.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.57.mlp.linear_fc2.weight |
| module.decoder.layers.58.mixer.in_proj.weight |
| module.decoder.layers.59.mlp.linear_fc2.weight |
| module.decoder.layers.56.mixer.in_proj.weight |
| module.decoder.layers.58.mixer.conv1d.weight |
| module.decoder.layers.57.mlp.linear_fc1.weight |
| module.decoder.layers.59.mlp.linear_fc1.weight |
| module.decoder.layers.58.mixer.out_proj.weight |
| module.decoder.layers.58.mixer.norm.weight |
| module.decoder.layers.56.mixer.out_proj.weight |
| module.decoder.layers.56.mixer.norm.weight |
| module.decoder.layers.56.mixer.conv1d.weight |
| Params for bucket 16 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.54.mixer.dt_bias |
| module.decoder.layers.56.mixer.dt_bias |
| module.decoder.layers.54.mixer.conv1d.bias |
| module.decoder.layers.52.mixer.conv1d.bias |
| module.decoder.layers.54.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.54.mixer.D |
| module.decoder.layers.56.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.56.mixer.A_log |
| module.decoder.layers.56.mixer.D |
| module.decoder.layers.55.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.54.mixer.A_log |
| module.decoder.layers.53.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.54.mixer.in_proj.weight |
| module.decoder.layers.55.mlp.linear_fc2.weight |
| module.decoder.layers.53.mlp.linear_fc2.weight |
| module.decoder.layers.52.mixer.in_proj.weight |
| module.decoder.layers.54.mixer.conv1d.weight |
| module.decoder.layers.53.mlp.linear_fc1.weight |
| module.decoder.layers.55.mlp.linear_fc1.weight |
| module.decoder.layers.54.mixer.out_proj.weight |
| module.decoder.layers.54.mixer.norm.weight |
| module.decoder.layers.52.mixer.out_proj.weight |
| module.decoder.layers.52.mixer.norm.weight |
| module.decoder.layers.52.mixer.conv1d.weight |
| Params for bucket 17 (41342874 elements, 41343232 padded size): |
| module.decoder.layers.50.mixer.dt_bias |
| module.decoder.layers.52.mixer.dt_bias |
| module.decoder.layers.49.mlp.linear_fc2.weight |
| module.decoder.layers.48.self_attention.linear_qkv.layer_norm_weight |
| module.decoder.layers.50.mixer.D |
| module.decoder.layers.52.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.52.mixer.A_log |
| module.decoder.layers.52.mixer.D |
| module.decoder.layers.51.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.50.mixer.norm.weight |
| module.decoder.layers.50.mixer.A_log |
| module.decoder.layers.48.self_attention.linear_qkv.weight |
| module.decoder.layers.50.mixer.in_proj.weight |
| module.decoder.layers.49.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.51.mlp.linear_fc2.weight |
| module.decoder.layers.50.mixer.conv1d.bias |
| module.decoder.layers.48.self_attention.linear_qkv.bias |
| module.decoder.layers.48.self_attention.linear_proj.weight |
| module.decoder.layers.50.mixer.conv1d.weight |
| module.decoder.layers.50.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.51.mlp.linear_fc1.weight |
| module.decoder.layers.50.mixer.out_proj.weight |
| module.decoder.layers.49.mlp.linear_fc1.weight |
| Params for bucket 18 (46174125 elements, 46174336 padded size): |
| module.decoder.layers.46.mixer.A_log |
| module.decoder.layers.45.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.47.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.46.mixer.D |
| module.decoder.layers.45.mlp.linear_fc2.weight |
| module.decoder.layers.44.mixer.norm.weight |
| module.decoder.layers.47.mlp.linear_fc2.weight |
| module.decoder.layers.46.mixer.norm.weight |
| module.decoder.layers.46.mixer.in_proj.weight |
| module.decoder.layers.44.mixer.in_proj.weight |
| module.decoder.layers.44.mixer.out_proj.weight |
| module.decoder.layers.45.mlp.linear_fc1.weight |
| module.decoder.layers.46.mixer.out_proj.weight |
| module.decoder.layers.47.mlp.linear_fc1.weight |
| module.decoder.layers.46.mixer.conv1d.weight |
| module.decoder.layers.46.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.44.mixer.conv1d.weight |
| module.decoder.layers.44.mixer.conv1d.bias |
| module.decoder.layers.46.mixer.conv1d.bias |
| module.decoder.layers.46.mixer.dt_bias |
| Params for bucket 19 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.42.mixer.A_log |
| module.decoder.layers.44.mixer.A_log |
| module.decoder.layers.44.mixer.D |
| module.decoder.layers.43.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.42.mixer.D |
| module.decoder.layers.41.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.40.mixer.norm.weight |
| module.decoder.layers.41.mlp.linear_fc2.weight |
| module.decoder.layers.43.mlp.linear_fc2.weight |
| module.decoder.layers.42.mixer.norm.weight |
| module.decoder.layers.42.mixer.in_proj.weight |
| module.decoder.layers.40.mixer.in_proj.weight |
| module.decoder.layers.42.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.42.mixer.conv1d.weight |
| module.decoder.layers.44.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.43.mlp.linear_fc1.weight |
| module.decoder.layers.42.mixer.out_proj.weight |
| module.decoder.layers.41.mlp.linear_fc1.weight |
| module.decoder.layers.40.mixer.out_proj.weight |
| module.decoder.layers.40.mixer.conv1d.bias |
| module.decoder.layers.40.mixer.conv1d.weight |
| module.decoder.layers.44.mixer.dt_bias |
| module.decoder.layers.42.mixer.conv1d.bias |
| module.decoder.layers.42.mixer.dt_bias |
| Params for bucket 20 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.38.mixer.conv1d.weight |
| module.decoder.layers.40.mixer.A_log |
| module.decoder.layers.38.mixer.out_proj.weight |
| module.decoder.layers.38.mixer.norm.weight |
| module.decoder.layers.37.mlp.linear_fc1.weight |
| module.decoder.layers.36.mixer.out_proj.weight |
| module.decoder.layers.36.mixer.norm.weight |
| module.decoder.layers.36.mixer.conv1d.weight |
| module.decoder.layers.40.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.39.mlp.linear_fc1.weight |
| module.decoder.layers.38.mixer.conv1d.bias |
| module.decoder.layers.38.mixer.dt_bias |
| module.decoder.layers.36.mixer.conv1d.bias |
| module.decoder.layers.38.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.37.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.38.mixer.A_log |
| module.decoder.layers.40.mixer.D |
| module.decoder.layers.39.mlp.linear_fc2.weight |
| module.decoder.layers.39.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.38.mixer.D |
| module.decoder.layers.37.mlp.linear_fc2.weight |
| module.decoder.layers.40.mixer.dt_bias |
| module.decoder.layers.38.mixer.in_proj.weight |
| module.decoder.layers.36.mixer.in_proj.weight |
| Params for bucket 21 (41342874 elements, 41343232 padded size): |
| module.decoder.layers.33.mlp.linear_fc1.weight |
| module.decoder.layers.34.mixer.conv1d.weight |
| module.decoder.layers.35.mlp.linear_fc1.weight |
| module.decoder.layers.34.mixer.out_proj.weight |
| module.decoder.layers.34.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.36.mixer.dt_bias |
| module.decoder.layers.34.mixer.dt_bias |
| module.decoder.layers.33.mlp.linear_fc2.weight |
| module.decoder.layers.32.self_attention.linear_qkv.layer_norm_weight |
| module.decoder.layers.34.mixer.D |
| module.decoder.layers.36.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.36.mixer.D |
| module.decoder.layers.36.mixer.A_log |
| module.decoder.layers.35.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.34.mixer.norm.weight |
| module.decoder.layers.34.mixer.A_log |
| module.decoder.layers.32.self_attention.linear_qkv.weight |
| module.decoder.layers.33.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.35.mlp.linear_fc2.weight |
| module.decoder.layers.34.mixer.conv1d.bias |
| module.decoder.layers.34.mixer.in_proj.weight |
| module.decoder.layers.32.self_attention.linear_qkv.bias |
| module.decoder.layers.32.self_attention.linear_proj.weight |
| Params for bucket 22 (46174125 elements, 46174336 padded size): |
| module.decoder.layers.30.mixer.dt_bias |
| module.decoder.layers.30.mixer.conv1d.bias |
| module.decoder.layers.28.mixer.conv1d.bias |
| module.decoder.layers.29.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.30.mixer.A_log |
| module.decoder.layers.31.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.30.mixer.D |
| module.decoder.layers.28.mixer.norm.weight |
| module.decoder.layers.31.mlp.linear_fc2.weight |
| module.decoder.layers.30.mixer.norm.weight |
| module.decoder.layers.30.mixer.in_proj.weight |
| module.decoder.layers.29.mlp.linear_fc2.weight |
| module.decoder.layers.28.mixer.in_proj.weight |
| module.decoder.layers.29.mlp.linear_fc1.weight |
| module.decoder.layers.28.mixer.out_proj.weight |
| module.decoder.layers.28.mixer.conv1d.weight |
| module.decoder.layers.30.mixer.out_proj.weight |
| module.decoder.layers.31.mlp.linear_fc1.weight |
| module.decoder.layers.30.mixer.conv1d.weight |
| module.decoder.layers.30.mixer.in_proj.layer_norm_weight |
| Params for bucket 23 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.26.mixer.dt_bias |
| module.decoder.layers.28.mixer.dt_bias |
| module.decoder.layers.26.mixer.conv1d.bias |
| module.decoder.layers.24.mixer.conv1d.bias |
| module.decoder.layers.25.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.26.mixer.D |
| module.decoder.layers.26.mixer.A_log |
| module.decoder.layers.28.mixer.D |
| module.decoder.layers.28.mixer.A_log |
| module.decoder.layers.27.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.27.mlp.linear_fc2.weight |
| module.decoder.layers.26.mixer.norm.weight |
| module.decoder.layers.25.mlp.linear_fc2.weight |
| module.decoder.layers.26.mixer.in_proj.weight |
| module.decoder.layers.24.mixer.norm.weight |
| module.decoder.layers.24.mixer.in_proj.weight |
| module.decoder.layers.25.mlp.linear_fc1.weight |
| module.decoder.layers.28.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.27.mlp.linear_fc1.weight |
| module.decoder.layers.26.mixer.out_proj.weight |
| module.decoder.layers.26.mixer.conv1d.weight |
| module.decoder.layers.26.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.24.mixer.out_proj.weight |
| module.decoder.layers.24.mixer.conv1d.weight |
| Params for bucket 24 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.22.mixer.dt_bias |
| module.decoder.layers.24.mixer.dt_bias |
| module.decoder.layers.22.mixer.conv1d.bias |
| module.decoder.layers.20.mixer.conv1d.bias |
| module.decoder.layers.21.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.22.mixer.D |
| module.decoder.layers.22.mixer.A_log |
| module.decoder.layers.24.mixer.D |
| module.decoder.layers.24.mixer.A_log |
| module.decoder.layers.23.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.23.mlp.linear_fc2.weight |
| module.decoder.layers.22.mixer.norm.weight |
| module.decoder.layers.21.mlp.linear_fc2.weight |
| module.decoder.layers.22.mixer.in_proj.weight |
| module.decoder.layers.20.mixer.norm.weight |
| module.decoder.layers.20.mixer.in_proj.weight |
| module.decoder.layers.21.mlp.linear_fc1.weight |
| module.decoder.layers.24.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.23.mlp.linear_fc1.weight |
| module.decoder.layers.22.mixer.out_proj.weight |
| module.decoder.layers.22.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.22.mixer.conv1d.weight |
| module.decoder.layers.20.mixer.out_proj.weight |
| module.decoder.layers.20.mixer.conv1d.weight |
| Params for bucket 25 (41342874 elements, 41343232 padded size): |
| module.decoder.layers.18.mixer.dt_bias |
| module.decoder.layers.20.mixer.dt_bias |
| module.decoder.layers.17.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.16.self_attention.linear_qkv.bias |
| module.decoder.layers.16.self_attention.linear_proj.weight |
| module.decoder.layers.18.mixer.A_log |
| module.decoder.layers.20.mixer.D |
| module.decoder.layers.20.mixer.A_log |
| module.decoder.layers.19.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.17.mlp.linear_fc1.weight |
| module.decoder.layers.18.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.17.mlp.linear_fc2.weight |
| module.decoder.layers.19.mlp.linear_fc2.weight |
| module.decoder.layers.18.mixer.conv1d.bias |
| module.decoder.layers.18.mixer.in_proj.weight |
| module.decoder.layers.16.self_attention.linear_qkv.layer_norm_weight |
| module.decoder.layers.18.mixer.D |
| module.decoder.layers.20.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.19.mlp.linear_fc1.weight |
| module.decoder.layers.18.mixer.out_proj.weight |
| module.decoder.layers.18.mixer.norm.weight |
| module.decoder.layers.18.mixer.conv1d.weight |
| module.decoder.layers.16.self_attention.linear_qkv.weight |
| Params for bucket 26 (46174125 elements, 46174336 padded size): |
| module.decoder.layers.13.mlp.linear_fc1.weight |
| module.decoder.layers.12.mixer.conv1d.weight |
| module.decoder.layers.15.mlp.linear_fc1.weight |
| module.decoder.layers.14.mixer.out_proj.weight |
| module.decoder.layers.14.mixer.conv1d.weight |
| module.decoder.layers.14.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.12.mixer.out_proj.weight |
| module.decoder.layers.14.mixer.conv1d.bias |
| module.decoder.layers.14.mixer.dt_bias |
| module.decoder.layers.12.mixer.conv1d.bias |
| module.decoder.layers.15.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.14.mixer.D |
| module.decoder.layers.14.mixer.A_log |
| module.decoder.layers.13.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.13.mlp.linear_fc2.weight |
| module.decoder.layers.12.mixer.norm.weight |
| module.decoder.layers.15.mlp.linear_fc2.weight |
| module.decoder.layers.14.mixer.norm.weight |
| module.decoder.layers.14.mixer.in_proj.weight |
| module.decoder.layers.12.mixer.in_proj.weight |
| Params for bucket 27 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.10.mixer.conv1d.weight |
| module.decoder.layers.9.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.12.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.11.mlp.linear_fc1.weight |
| module.decoder.layers.10.mixer.out_proj.weight |
| module.decoder.layers.10.mixer.norm.weight |
| module.decoder.layers.9.mlp.linear_fc2.weight |
| module.decoder.layers.12.mixer.dt_bias |
| module.decoder.layers.10.mixer.dt_bias |
| module.decoder.layers.8.mixer.norm.weight |
| module.decoder.layers.8.mixer.in_proj.weight |
| module.decoder.layers.9.mlp.linear_fc1.weight |
| module.decoder.layers.12.mixer.A_log |
| module.decoder.layers.12.mixer.D |
| module.decoder.layers.10.mixer.conv1d.bias |
| module.decoder.layers.10.mixer.A_log |
| module.decoder.layers.10.mixer.D |
| module.decoder.layers.8.mixer.out_proj.weight |
| module.decoder.layers.8.mixer.conv1d.weight |
| module.decoder.layers.10.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.10.mixer.in_proj.weight |
| module.decoder.layers.11.mlp.linear_fc2.weight |
| module.decoder.layers.11.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.8.mixer.conv1d.bias |
| Params for bucket 28 (46176090 elements, 46176384 padded size): |
| module.decoder.layers.5.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.8.mixer.D |
| module.decoder.layers.8.mixer.A_log |
| module.decoder.layers.7.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.6.mixer.D |
| module.decoder.layers.6.mixer.A_log |
| module.decoder.layers.6.mixer.in_proj.weight |
| module.decoder.layers.5.mlp.linear_fc2.weight |
| module.decoder.layers.7.mlp.linear_fc2.weight |
| module.decoder.layers.6.mixer.norm.weight |
| module.decoder.layers.4.mixer.conv1d.bias |
| module.decoder.layers.4.mixer.in_proj.weight |
| module.decoder.layers.6.mixer.conv1d.weight |
| module.decoder.layers.5.mlp.linear_fc1.weight |
| module.decoder.layers.8.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.7.mlp.linear_fc1.weight |
| module.decoder.layers.6.mixer.out_proj.weight |
| module.decoder.layers.6.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.4.mixer.out_proj.weight |
| module.decoder.layers.4.mixer.norm.weight |
| module.decoder.layers.4.mixer.conv1d.weight |
| module.decoder.layers.6.mixer.dt_bias |
| module.decoder.layers.8.mixer.dt_bias |
| module.decoder.layers.6.mixer.conv1d.bias |
| Params for bucket 29 (41342874 elements, 41343232 padded size): |
| module.decoder.layers.2.mixer.dt_bias |
| module.decoder.layers.4.mixer.A_log |
| module.decoder.layers.3.mlp.linear_fc1.weight |
| module.decoder.layers.0.self_attention.linear_qkv.weight |
| module.decoder.layers.0.self_attention.linear_proj.weight |
| module.decoder.layers.4.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.3.mlp.linear_fc1.layer_norm_weight |
| module.decoder.layers.2.mixer.A_log |
| module.decoder.layers.0.self_attention.linear_qkv.bias |
| module.decoder.layers.2.mixer.in_proj.weight |
| module.decoder.layers.1.mlp.linear_fc2.weight |
| module.decoder.layers.2.mixer.D |
| module.decoder.layers.3.mlp.linear_fc2.weight |
| module.decoder.layers.2.mixer.norm.weight |
| module.decoder.layers.2.mixer.conv1d.bias |
| module.decoder.layers.0.self_attention.linear_qkv.layer_norm_weight |
| module.decoder.layers.2.mixer.conv1d.weight |
| module.decoder.layers.1.mlp.linear_fc1.weight |
| module.decoder.layers.4.mixer.dt_bias |
| module.decoder.layers.4.mixer.D |
| module.decoder.layers.2.mixer.out_proj.weight |
| module.decoder.layers.2.mixer.in_proj.layer_norm_weight |
| module.decoder.layers.1.mlp.linear_fc1.layer_norm_weight |
| Params for bucket 30 (95109120 elements, 95109120 padded size): |
| module.embedding.word_embeddings.weight |
| [MambaModel( |
| (embedding): LanguageModelEmbedding( |
| (word_embeddings): VocabParallelEmbedding() |
| (embedding_dropout): Dropout(p=0.0, inplace=False) |
| ) |
| (rotary_pos_emb): RotaryEmbedding() |
| (decoder): MambaStack( |
| (layers): ModuleList( |
| (0): TransformerLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): SelfAttention( |
| (core_attention): TEDotProductAttention( |
| (flash_attention): FlashAttention() |
| (fused_attention): FusedAttention() |
| (unfused_attention): UnfusedDotProductAttention( |
| (scale_mask_softmax): FusedScaleMaskSoftmax() |
| (attention_dropout): Dropout(p=0.0, inplace=False) |
| ) |
| ) |
| (linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| (linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2) |
| (q_layernorm): IdentityOp() |
| (k_layernorm): IdentityOp() |
| ) |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): IdentityOp() |
| (mlp_bda): IdentityFuncOp() |
| ) |
| (1): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (2): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (3): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (4): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (5): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (6): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (7): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (8): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (9): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (10): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (11): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (12): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (13): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (14): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (15): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (16): TransformerLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): SelfAttention( |
| (core_attention): TEDotProductAttention( |
| (flash_attention): FlashAttention() |
| (fused_attention): FusedAttention() |
| (unfused_attention): UnfusedDotProductAttention( |
| (scale_mask_softmax): FusedScaleMaskSoftmax() |
| (attention_dropout): Dropout(p=0.0, inplace=False) |
| ) |
| ) |
| (linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| (linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2) |
| (q_layernorm): IdentityOp() |
| (k_layernorm): IdentityOp() |
| ) |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): IdentityOp() |
| (mlp_bda): IdentityFuncOp() |
| ) |
| (17): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (18): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (19): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (20): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (21): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (22): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (23): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (24): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (25): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (26): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (27): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (28): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (29): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (30): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (31): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (32): TransformerLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): SelfAttention( |
| (core_attention): TEDotProductAttention( |
| (flash_attention): FlashAttention() |
| (fused_attention): FusedAttention() |
| (unfused_attention): UnfusedDotProductAttention( |
| (scale_mask_softmax): FusedScaleMaskSoftmax() |
| (attention_dropout): Dropout(p=0.0, inplace=False) |
| ) |
| ) |
| (linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| (linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2) |
| (q_layernorm): IdentityOp() |
| (k_layernorm): IdentityOp() |
| ) |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): IdentityOp() |
| (mlp_bda): IdentityFuncOp() |
| ) |
| (33): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (34): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (35): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (36): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (37): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (38): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (39): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (40): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (41): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (42): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (43): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (44): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (45): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (46): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (47): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (48): TransformerLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): SelfAttention( |
| (core_attention): TEDotProductAttention( |
| (flash_attention): FlashAttention() |
| (fused_attention): FusedAttention() |
| (unfused_attention): UnfusedDotProductAttention( |
| (scale_mask_softmax): FusedScaleMaskSoftmax() |
| (attention_dropout): Dropout(p=0.0, inplace=False) |
| ) |
| ) |
| (linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| (linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2) |
| (q_layernorm): IdentityOp() |
| (k_layernorm): IdentityOp() |
| ) |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): IdentityOp() |
| (mlp_bda): IdentityFuncOp() |
| ) |
| (49): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (50): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (51): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (52): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (53): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (54): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (55): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (56): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (57): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (58): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (59): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (60): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (61): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (62): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (63): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (64): TransformerLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): SelfAttention( |
| (core_attention): TEDotProductAttention( |
| (flash_attention): FlashAttention() |
| (fused_attention): FusedAttention() |
| (unfused_attention): UnfusedDotProductAttention( |
| (scale_mask_softmax): FusedScaleMaskSoftmax() |
| (attention_dropout): Dropout(p=0.0, inplace=False) |
| ) |
| ) |
| (linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| (linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2) |
| (q_layernorm): IdentityOp() |
| (k_layernorm): IdentityOp() |
| ) |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): IdentityOp() |
| (mlp_bda): IdentityFuncOp() |
| ) |
| (65): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (66): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (67): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (68): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (69): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (70): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (71): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (72): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (73): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (74): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (75): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (76): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (77): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (78): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (79): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (80): TransformerLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): SelfAttention( |
| (core_attention): TEDotProductAttention( |
| (flash_attention): FlashAttention() |
| (fused_attention): FusedAttention() |
| (unfused_attention): UnfusedDotProductAttention( |
| (scale_mask_softmax): FusedScaleMaskSoftmax() |
| (attention_dropout): Dropout(p=0.0, inplace=False) |
| ) |
| ) |
| (linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| (linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2) |
| (q_layernorm): IdentityOp() |
| (k_layernorm): IdentityOp() |
| ) |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): IdentityOp() |
| (mlp_bda): IdentityFuncOp() |
| ) |
| (81): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (82): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (83): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (84): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (85): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (86): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (87): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (88): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (89): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (90): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (91): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (92): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (93): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (94): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (95): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (96): TransformerLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): SelfAttention( |
| (core_attention): TEDotProductAttention( |
| (flash_attention): FlashAttention() |
| (fused_attention): FusedAttention() |
| (unfused_attention): UnfusedDotProductAttention( |
| (scale_mask_softmax): FusedScaleMaskSoftmax() |
| (attention_dropout): Dropout(p=0.0, inplace=False) |
| ) |
| ) |
| (linear_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| (linear_qkv): TELayerNormColumnParallelLinear(in_features=1920, out_features=1344, bias=True, TP=2) |
| (q_layernorm): IdentityOp() |
| (k_layernorm): IdentityOp() |
| ) |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): IdentityOp() |
| (mlp_bda): IdentityFuncOp() |
| ) |
| (97): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (98): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (99): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (100): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (101): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (102): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (103): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (104): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (105): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (106): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (107): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (108): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (109): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| (110): MambaLayer( |
| (mixer): MambaMixer( |
| (in_proj): TELayerNormColumnParallelLinear(in_features=1920, out_features=3855, bias=False, TP=2) |
| (conv1d): Conv1d(2880, 2880, kernel_size=(4,), stride=(1,), padding=(3,), groups=2880) |
| (act): SiLU() |
| (norm): ExtendedRMSNorm() |
| (out_proj): TERowParallelLinear(in_features=960, out_features=1920, bias=False, TP=2) |
| ) |
| (norm): IdentityOp() |
| ) |
| (111): MLPLayer( |
| (input_layernorm): IdentityOp() |
| (self_attention): IdentityOp() |
| (self_attn_bda): IdentityFuncOp() |
| (pre_cross_attn_layernorm): IdentityOp() |
| (cross_attention): IdentityOp() |
| (cross_attn_bda): IdentityFuncOp() |
| (pre_mlp_layernorm): IdentityOp() |
| (mlp): MLP( |
| (linear_fc1): TELayerNormColumnParallelLinear(in_features=1920, out_features=4800, bias=False, TP=2) |
| (linear_fc2): TERowParallelLinear(in_features=2400, out_features=1920, bias=False, TP=2) |
| ) |
| ) |
| ) |
| (final_norm): RMSNorm() |
| ) |
| (output_layer): ColumnParallelLinear(in_features=1920, out_features=99072, bias=False, TP=2) |
| )] |
| [DEBUG] freeze_non_mamba: False |
| INFO:megatron.core.optimizer:Setting up optimizer with config OptimizerConfig(optimizer='adam', lr=2e-05, min_lr=7e-07, decoupled_lr=None, decoupled_min_lr=None, weight_decay=0.1, fp16=False, bf16=True, params_dtype=torch.bfloat16, use_precision_aware_optimizer=False, main_grads_dtype=torch.float32, main_params_dtype=torch.float32, exp_avg_dtype=torch.float32, exp_avg_sq_dtype=torch.float32, loss_scale=None, initial_loss_scale=4294967296, min_loss_scale=1.0, loss_scale_window=1000, hysteresis=2, adam_beta1=0.9, adam_beta2=0.999, adam_eps=1e-08, sgd_momentum=0.9, muon_momentum=0.95, muon_nesterov=True, muon_ns_steps=5, muon_matched_adamw_rms=0.2, use_distributed_optimizer=True, overlap_param_gather_with_optimizer_step=False, optimizer_cpu_offload=False, optimizer_offload_fraction=1.0, use_torch_optimizer_for_cpu_offload=False, overlap_cpu_optimizer_d2h_h2d=False, pin_cpu_grads=True, pin_cpu_params=True, clip_grad=0.5, log_num_zeros_in_grad=False, barrier_with_L1_time=True, timers=<megatron.core.timers.Timers object at 0x7efd9ec4c350>, config_logger_dir='') |
| setting training iterations to 59 |
| INFO:megatron.core.optimizer_param_scheduler:> learning rate decay style: linear |
| [DEBUG] freeze_non_mamba: False |
| [DEBUG] freeze_non_mamba: False |
| [DEBUG] freeze_non_mamba: False |
| [DEBUG] freeze_non_mamba: False |
| loading checkpoint from /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/RADLADS-paper/out/L56-D1920-qwen_mamba2_qwen2-e1-i1920-s320-hd64-gn6-A0-S512--step1-dclm10b/rwkv-394-hf-A7-0_8_16_24_32_40_48/megatron-pp1-tp2 at iteration 0 |
| could not find arguments in the checkpoint ... |
| checkpoint version 0 |
| successfully fixed query-key-values ordering for checkpoint version 0 |
| successfully loaded checkpoint from /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/RADLADS-paper/out/L56-D1920-qwen_mamba2_qwen2-e1-i1920-s320-hd64-gn6-A0-S512--step1-dclm10b/rwkv-394-hf-A7-0_8_16_24_32_40_48/megatron-pp1-tp2 [ t 1/2, p 1/1 ] at iteration 0 |
| (min, max) time across ranks (ms): |
| load-checkpoint ................................: (6511.67, 6511.73) |
| [after model, optimizer, and learning rate scheduler are built] datetime: 2025-09-17 22:26:16 |
| saving checkpoint at iteration 0 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 in torch format |
| successfully saved checkpoint from iteration 0 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 [ t 1/2, p 1/1 ] |
| > building train, validation, and test datasets ... |
| > datasets target sizes (minimum size): |
| train: 61035 |
| validation: 10240 |
| test: 10240 |
| INFO:megatron.core.datasets.blended_megatron_dataset_config:Let split_matrix = [(0, 1.0), None, None] |
| > building train, validation, and test datasets for GPT ... |
| INFO:megatron.core.datasets.blended_megatron_dataset_builder:Building GPTDataset splits with sizes=(61035, 10240, 10240) and config=GPTDatasetConfig(random_seed=1234, sequence_length=32768, blend=(['/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache/datasets/huggingface/Teaven/combine_2B_0908/binidx/yulan_mini'], None), blend_per_split=None, split='100,0,0', split_matrix=[(0, 1.0), None, None], num_dataset_builder_threads=1, path_to_cache='/mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache', mmap_bin_files=True, mock=False, tokenizer=<megatron.training.tokenizer.tokenizer._HuggingFaceTokenizer object at 0x7efe4d2af6d0>, mid_level_dataset_surplus=0.005, reset_position_ids=False, reset_attention_mask=False, eod_mask_loss=False, create_attention_mask=False, drop_last_partial_validation_sequence=True, add_extra_token_to_sequence=True, object_storage_cache_path=None) |
| INFO:megatron.core.datasets.indexed_dataset:Load the _IndexReader from /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/cache/datasets/huggingface/Teaven/combine_2B_0908/binidx/yulan_mini.idx |
| INFO:megatron.core.datasets.indexed_dataset: Extract the sequence lengths |
| INFO:megatron.core.datasets.indexed_dataset: Extract the sequence pointers |
| INFO:megatron.core.datasets.indexed_dataset: Extract the document indices |
| INFO:megatron.core.datasets.indexed_dataset:> total number of sequences: 73736 |
| INFO:megatron.core.datasets.indexed_dataset:> total number of documents: 73736 |
| INFO:megatron.core.datasets.gpt_dataset:Load the GPTDataset train indices |
| INFO:megatron.core.datasets.gpt_dataset: Load the document index from 195a986bb6005248efb3de26e9334e88-GPTDataset-train-document_index.npy |
| INFO:megatron.core.datasets.gpt_dataset: Load the sample index from 195a986bb6005248efb3de26e9334e88-GPTDataset-train-sample_index.npy |
| INFO:megatron.core.datasets.gpt_dataset: Load the shuffle index from 195a986bb6005248efb3de26e9334e88-GPTDataset-train-shuffle_index.npy |
| INFO:megatron.core.datasets.gpt_dataset:> total number of samples: 64518 |
| > finished creating GPT datasets ... |
| [after dataloaders are built] datetime: 2025-09-17 22:26:22 |
| done with setup ... |
| (min, max) time across ranks (ms): |
| model-and-optimizer-setup ......................: (8145.52, 8201.88) |
| train/valid/test-data-iterators-setup ..........: (12.45, 100.39) |
| training ... |
| Setting rerun_state_machine.current_iteration to 0... |
| [before the start of training step] datetime: 2025-09-17 22:26:22 |
| Iter 0 increase timing-log-level to 2. |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| [Rank 2] (after 1 iterations) memory (MB) | allocated: 12648.54443359375 | max allocated: 44991.1953125 | reserved: 47306.0 | max reserved: 47306.0 | MEM: 48.63% |
| [Rank 3] (after 1 iterations) memory (MB) | allocated: 12648.54443359375 | max allocated: 44991.1953125 | reserved: 47306.0 | max reserved: 47306.0 | MEM: 48.63% |
| [Rank 1] (after 1 iterations) memory (MB) | allocated: 12648.5341796875 | max allocated: 44991.19189453125 | reserved: 47224.0 | max reserved: 47224.0 | MEM: 48.54% |
| [2025-09-17 22:44:07] iteration 1/ 59 | consumed samples: 1024 | elapsed time per iteration (ms): 1065339.4 | throughput per GPU (TFLOP/s/GPU): 267.5 | MFU 27.05% | learning rate: 6.712553E-06 | global batch size: 1024 | lm loss: 3.847867E+00 | loss scale: 1.0 | grad norm: 581559296.000 | num zeros: 50149872.0 | params norm: 9733.412 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 17:09:49.684691 | finish at 2025-09-18 15:54:08 |
| Number of parameters in transformer block in billions: 4.09 |
| Number of parameters in embedding layers in billions: 0.38 |
| Total number of parameters in billions: 4.47 |
| Number of parameters in most loaded shard in billions: 2.2344 |
| Activation memory footprint per transformer layer: 840.0 MB |
| Theoretical memory footprints: weight and optimizer=25570.57 MB, activation=100422.12 MB, total=125992.69 MB |
|
|
| [Rank 0] (after 1 iterations) memory (MB) | allocated: 12648.5341796875 | max allocated: 44991.19189453125 | reserved: 47224.0 | max reserved: 47224.0 | MEM: 48.54% |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/YuLan-Pretrain/megatron/core/pipeline_parallel/schedules.py:157: UserWarning: c10d::allreduce_: an autograd kernel was not registered to the Autograd key(s) but we are trying to backprop through it. This may lead to silently incorrect behavior. This behavior is deprecated and will be removed in a future version of PyTorch. If your operator is differentiable, please ensure you have registered an autograd kernel to the correct Autograd key (e.g. DispatchKey::Autograd, DispatchKey::CompositeImplicitAutograd). If your operator is not differentiable, or to squash this warning and use the previous behavior, please register torch::CppFunction::makeFallthrough() to DispatchKey::Autograd. (Triggered internally at /pytorch/torch/csrc/autograd/autograd_not_implemented_fallback.cpp:62.) |
| Variable._execution_engine.run_backward( |
| [2025-09-17 23:01:03] iteration 2/ 59 | consumed samples: 2048 | elapsed time per iteration (ms): 1016177.4 | throughput per GPU (TFLOP/s/GPU): 280.4 | MFU 28.36% | learning rate: 1.342511E-05 | global batch size: 1024 | lm loss: 3.936432E+00 | loss scale: 1.0 | grad norm: 2996916736.000 | num zeros: 49567872.0 | params norm: 9733.406 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 16:05:22.111929 | finish at 2025-09-18 15:06:25 |
| [2025-09-17 23:17:48] iteration 3/ 59 | consumed samples: 3072 | elapsed time per iteration (ms): 1005115.4 | throughput per GPU (TFLOP/s/GPU): 283.5 | MFU 28.67% | learning rate: 1.999301E-05 | global batch size: 1024 | lm loss: 3.885512E+00 | loss scale: 1.0 | grad norm: 3171558400.000 | num zeros: 48767412.0 | params norm: 9733.392 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 15:38:06.463900 | finish at 2025-09-18 14:55:55 |
| [2025-09-17 23:34:34] iteration 4/ 59 | consumed samples: 4096 | elapsed time per iteration (ms): 1005725.7 | throughput per GPU (TFLOP/s/GPU): 283.4 | MFU 28.65% | learning rate: 1.965217E-05 | global batch size: 1024 | lm loss: 3.872411E+00 | loss scale: 1.0 | grad norm: 28114848.000 | num zeros: 49829060.0 | params norm: 9733.373 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 15:21:54.915160 | finish at 2025-09-18 14:56:29 |
| [2025-09-17 23:51:19] iteration 5/ 59 | consumed samples: 5120 | elapsed time per iteration (ms): 1004912.5 | throughput per GPU (TFLOP/s/GPU): 283.6 | MFU 28.67% | learning rate: 1.931133E-05 | global batch size: 1024 | lm loss: 3.757929E+00 | loss scale: 1.0 | grad norm: 1324076160.000 | num zeros: 49525848.0 | params norm: 9733.354 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 15:04:25.272742 | finish at 2025-09-18 14:55:44 |
| Iter 5 reset timing-log-level. |
| [2025-09-18 00:08:00] iteration 6/ 59 | consumed samples: 6144 | elapsed time per iteration (ms): 1000789.1 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.897049E-05 | global batch size: 1024 | lm loss: 3.720438E+00 | loss scale: 1.0 | grad norm: 128727096.000 | num zeros: 48811584.0 | params norm: 9733.335 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 14:44:01.821778 | finish at 2025-09-18 14:52:01 |
| [2025-09-18 00:24:41] iteration 7/ 59 | consumed samples: 7168 | elapsed time per iteration (ms): 1000967.1 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.862965E-05 | global batch size: 1024 | lm loss: 3.737079E+00 | loss scale: 1.0 | grad norm: 2022209152.000 | num zeros: 47870832.0 | params norm: 9733.317 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 14:27:30.288079 | finish at 2025-09-18 14:52:11 |
| [2025-09-18 00:41:21] iteration 8/ 59 | consumed samples: 8192 | elapsed time per iteration (ms): 1000824.7 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.828882E-05 | global batch size: 1024 | lm loss: 3.762561E+00 | loss scale: 1.0 | grad norm: 5132039680.000 | num zeros: 49402816.0 | params norm: 9733.299 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 14:10:42.058004 | finish at 2025-09-18 14:52:03 |
| [2025-09-18 00:58:02] iteration 9/ 59 | consumed samples: 9216 | elapsed time per iteration (ms): 1000899.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.794798E-05 | global batch size: 1024 | lm loss: 3.833430E+00 | loss scale: 1.0 | grad norm: 8436005888.000 | num zeros: 49748140.0 | params norm: 9733.281 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 13:54:04.963479 | finish at 2025-09-18 14:52:07 |
| [2025-09-18 01:14:43] iteration 10/ 59 | consumed samples: 10240 | elapsed time per iteration (ms): 1000818.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.760714E-05 | global batch size: 1024 | lm loss: 3.975019E+00 | loss scale: 1.0 | grad norm: 13745026048.000 | num zeros: 48563464.0 | params norm: 9733.263 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 13:37:20.092904 | finish at 2025-09-18 14:52:03 |
| [2025-09-18 01:31:24] iteration 11/ 59 | consumed samples: 11264 | elapsed time per iteration (ms): 1000845.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.726630E-05 | global batch size: 1024 | lm loss: 4.137877E+00 | loss scale: 1.0 | grad norm: 110488526848.000 | num zeros: 50133848.0 | params norm: 9733.247 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 13:20:40.568092 | finish at 2025-09-18 14:52:05 |
| [2025-09-18 01:48:05] iteration 12/ 59 | consumed samples: 12288 | elapsed time per iteration (ms): 1000906.4 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.692546E-05 | global batch size: 1024 | lm loss: 4.174570E+00 | loss scale: 1.0 | grad norm: 213260451840.000 | num zeros: 49354312.0 | params norm: 9733.230 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 13:04:02.599566 | finish at 2025-09-18 14:52:07 |
| [2025-09-18 02:04:46] iteration 13/ 59 | consumed samples: 13312 | elapsed time per iteration (ms): 1000826.5 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.658462E-05 | global batch size: 1024 | lm loss: 4.318902E+00 | loss scale: 1.0 | grad norm: 29837592576.000 | num zeros: 49300288.0 | params norm: 9733.213 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 12:47:18.019107 | finish at 2025-09-18 14:52:04 |
| [2025-09-18 02:21:26] iteration 14/ 59 | consumed samples: 14336 | elapsed time per iteration (ms): 1000652.2 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.80% | learning rate: 1.624378E-05 | global batch size: 1024 | lm loss: 4.142406E+00 | loss scale: 1.0 | grad norm: 5777787904.000 | num zeros: 47241792.0 | params norm: 9733.197 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 12:30:29.348720 | finish at 2025-09-18 14:51:56 |
| [2025-09-18 02:38:07] iteration 15/ 59 | consumed samples: 15360 | elapsed time per iteration (ms): 1000613.6 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.80% | learning rate: 1.590294E-05 | global batch size: 1024 | lm loss: 4.086899E+00 | loss scale: 1.0 | grad norm: 48051974144.000 | num zeros: 49042056.0 | params norm: 9733.181 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 12:13:46.997184 | finish at 2025-09-18 14:51:54 |
| [2025-09-18 02:54:48] iteration 16/ 59 | consumed samples: 16384 | elapsed time per iteration (ms): 1000891.7 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.556211E-05 | global batch size: 1024 | lm loss: 3.986907E+00 | loss scale: 1.0 | grad norm: 1339338496.000 | num zeros: 49090020.0 | params norm: 9733.166 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 11:57:18.344998 | finish at 2025-09-18 14:52:06 |
| [2025-09-18 03:11:29] iteration 17/ 59 | consumed samples: 17408 | elapsed time per iteration (ms): 1000751.6 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.522127E-05 | global batch size: 1024 | lm loss: 3.926814E+00 | loss scale: 1.0 | grad norm: 751090048.000 | num zeros: 49013412.0 | params norm: 9733.150 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 11:40:31.568773 | finish at 2025-09-18 14:52:00 |
| [2025-09-18 03:28:09] iteration 18/ 59 | consumed samples: 18432 | elapsed time per iteration (ms): 1000766.6 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.488043E-05 | global batch size: 1024 | lm loss: 3.843276E+00 | loss scale: 1.0 | grad norm: 4079015936.000 | num zeros: 48842440.0 | params norm: 9733.136 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 11:23:51.429882 | finish at 2025-09-18 14:52:01 |
| [2025-09-18 03:44:50] iteration 19/ 59 | consumed samples: 19456 | elapsed time per iteration (ms): 1000849.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.453959E-05 | global batch size: 1024 | lm loss: 3.819354E+00 | loss scale: 1.0 | grad norm: 663727040.000 | num zeros: 48992272.0 | params norm: 9733.121 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 11:07:13.967733 | finish at 2025-09-18 14:52:04 |
| [2025-09-18 04:01:31] iteration 20/ 59 | consumed samples: 20480 | elapsed time per iteration (ms): 1000851.7 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.419875E-05 | global batch size: 1024 | lm loss: 3.898749E+00 | loss scale: 1.0 | grad norm: 177677072.000 | num zeros: 49040544.0 | params norm: 9733.107 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 10:50:33.214592 | finish at 2025-09-18 14:52:04 |
| [2025-09-18 04:18:12] iteration 21/ 59 | consumed samples: 21504 | elapsed time per iteration (ms): 1000741.4 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.385791E-05 | global batch size: 1024 | lm loss: 4.066440E+00 | loss scale: 1.0 | grad norm: 102174968.000 | num zeros: 49475984.0 | params norm: 9733.093 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 10:33:48.171461 | finish at 2025-09-18 14:52:00 |
| [2025-09-18 04:34:53] iteration 22/ 59 | consumed samples: 22528 | elapsed time per iteration (ms): 1000731.2 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.351707E-05 | global batch size: 1024 | lm loss: 4.138809E+00 | loss scale: 1.0 | grad norm: 1939881984.000 | num zeros: 49917468.0 | params norm: 9733.080 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 10:17:07.053182 | finish at 2025-09-18 14:52:00 |
| [2025-09-18 04:51:33] iteration 23/ 59 | consumed samples: 23552 | elapsed time per iteration (ms): 1000630.4 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.80% | learning rate: 1.317623E-05 | global batch size: 1024 | lm loss: 4.354475E+00 | loss scale: 1.0 | grad norm: 1467312896.000 | num zeros: 49097692.0 | params norm: 9733.067 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 10:00:22.694501 | finish at 2025-09-18 14:51:56 |
| [2025-09-18 05:08:14] iteration 24/ 59 | consumed samples: 24576 | elapsed time per iteration (ms): 1000795.8 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.283539E-05 | global batch size: 1024 | lm loss: 4.601981E+00 | loss scale: 1.0 | grad norm: 23645913088.000 | num zeros: 49239816.0 | params norm: 9733.054 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 9:43:47.853149 | finish at 2025-09-18 14:52:02 |
| [2025-09-18 05:24:55] iteration 25/ 59 | consumed samples: 25600 | elapsed time per iteration (ms): 1000910.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.249456E-05 | global batch size: 1024 | lm loss: 4.786105E+00 | loss scale: 1.0 | grad norm: 12307973120.000 | num zeros: 49598280.0 | params norm: 9733.041 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 9:27:10.949439 | finish at 2025-09-18 14:52:06 |
| [2025-09-18 05:41:36] iteration 26/ 59 | consumed samples: 26624 | elapsed time per iteration (ms): 1000848.6 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.215372E-05 | global batch size: 1024 | lm loss: 5.074014E+00 | loss scale: 1.0 | grad norm: 26155917312.000 | num zeros: 49150844.0 | params norm: 9733.029 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 9:10:28.003309 | finish at 2025-09-18 14:52:04 |
| [2025-09-18 05:58:17] iteration 27/ 59 | consumed samples: 27648 | elapsed time per iteration (ms): 1000920.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.181288E-05 | global batch size: 1024 | lm loss: 5.265514E+00 | loss scale: 1.0 | grad norm: 124921552896.000 | num zeros: 49020472.0 | params norm: 9733.018 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 8:53:49.449287 | finish at 2025-09-18 14:52:06 |
| [2025-09-18 06:14:58] iteration 28/ 59 | consumed samples: 28672 | elapsed time per iteration (ms): 1000958.1 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.147204E-05 | global batch size: 1024 | lm loss: 5.399610E+00 | loss scale: 1.0 | grad norm: 7730917888.000 | num zeros: 49531044.0 | params norm: 9733.006 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 8:37:09.699772 | finish at 2025-09-18 14:52:07 |
| [2025-09-18 06:31:38] iteration 29/ 59 | consumed samples: 29696 | elapsed time per iteration (ms): 1000807.7 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 1.113120E-05 | global batch size: 1024 | lm loss: 5.464283E+00 | loss scale: 1.0 | grad norm: 1776406757376.000 | num zeros: 49643804.0 | params norm: 9732.995 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 8:20:24.230998 | finish at 2025-09-18 14:52:03 |
| saving checkpoint at iteration 29 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 in torch format |
| successfully saved checkpoint from iteration 29 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 [ t 1/2, p 1/1 ] |
| (min, max) time across ranks (ms): |
| save-checkpoint ................................: (51941.34, 51941.37) |
| [2025-09-18 06:49:11] iteration 30/ 59 | consumed samples: 30720 | elapsed time per iteration (ms): 1001022.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.079036E-05 | global batch size: 1024 | lm loss: 5.466630E+00 | loss scale: 1.0 | grad norm: 201713991680.000 | num zeros: 48573168.0 | params norm: 9732.984 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 8:03:49.644004 | finish at 2025-09-18 14:53:01 |
| [2025-09-18 07:05:52] iteration 31/ 59 | consumed samples: 31744 | elapsed time per iteration (ms): 1000966.6 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.044952E-05 | global batch size: 1024 | lm loss: 5.530435E+00 | loss scale: 1.0 | grad norm: 16827718656.000 | num zeros: 49106456.0 | params norm: 9732.973 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 7:47:07.063577 | finish at 2025-09-18 14:52:59 |
| [2025-09-18 07:22:33] iteration 32/ 59 | consumed samples: 32768 | elapsed time per iteration (ms): 1001076.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 1.010868E-05 | global batch size: 1024 | lm loss: 5.704006E+00 | loss scale: 1.0 | grad norm: 279325507584.000 | num zeros: 48570972.0 | params norm: 9732.963 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 7:30:29.056911 | finish at 2025-09-18 14:53:02 |
| [2025-09-18 07:39:14] iteration 33/ 59 | consumed samples: 33792 | elapsed time per iteration (ms): 1000973.1 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 9.767845E-06 | global batch size: 1024 | lm loss: 5.747099E+00 | loss scale: 1.0 | grad norm: 31260119040.000 | num zeros: 49229308.0 | params norm: 9732.954 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 7:13:45.300946 | finish at 2025-09-18 14:53:00 |
| [2025-09-18 07:55:55] iteration 34/ 59 | consumed samples: 34816 | elapsed time per iteration (ms): 1000882.6 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 9.427005E-06 | global batch size: 1024 | lm loss: 5.803835E+00 | loss scale: 1.0 | grad norm: 640638255104.000 | num zeros: 49311708.0 | params norm: 9732.944 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 6:57:02.066164 | finish at 2025-09-18 14:52:57 |
| [2025-09-18 08:12:36] iteration 35/ 59 | consumed samples: 35840 | elapsed time per iteration (ms): 1000906.9 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 9.086167E-06 | global batch size: 1024 | lm loss: 5.908048E+00 | loss scale: 1.0 | grad norm: 101146836992.000 | num zeros: 48868296.0 | params norm: 9732.935 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 6:40:21.766422 | finish at 2025-09-18 14:52:58 |
| [2025-09-18 08:29:17] iteration 36/ 59 | consumed samples: 36864 | elapsed time per iteration (ms): 1000991.8 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 8.745328E-06 | global batch size: 1024 | lm loss: 6.095221E+00 | loss scale: 1.0 | grad norm: 461121159168.000 | num zeros: 48946824.0 | params norm: 9732.926 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 6:23:42.811659 | finish at 2025-09-18 14:53:00 |
| [2025-09-18 08:45:58] iteration 37/ 59 | consumed samples: 37888 | elapsed time per iteration (ms): 1000857.1 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 8.404490E-06 | global batch size: 1024 | lm loss: 5.988223E+00 | loss scale: 1.0 | grad norm: 50295472128.000 | num zeros: 49686208.0 | params norm: 9732.918 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 6:06:58.856651 | finish at 2025-09-18 14:52:57 |
| [2025-09-18 09:02:39] iteration 38/ 59 | consumed samples: 38912 | elapsed time per iteration (ms): 1001077.2 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 8.063650E-06 | global batch size: 1024 | lm loss: 5.932316E+00 | loss scale: 1.0 | grad norm: 465845452800.000 | num zeros: 49469216.0 | params norm: 9732.909 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 5:50:22.621344 | finish at 2025-09-18 14:53:02 |
| [2025-09-18 09:19:20] iteration 39/ 59 | consumed samples: 39936 | elapsed time per iteration (ms): 1000904.9 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 7.722811E-06 | global batch size: 1024 | lm loss: 5.991258E+00 | loss scale: 1.0 | grad norm: 100241760256.000 | num zeros: 49582496.0 | params norm: 9732.901 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 5:33:38.097820 | finish at 2025-09-18 14:52:58 |
| [2025-09-18 09:36:01] iteration 40/ 59 | consumed samples: 40960 | elapsed time per iteration (ms): 1001159.5 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 7.381972E-06 | global batch size: 1024 | lm loss: 5.912158E+00 | loss scale: 1.0 | grad norm: 39360831488.000 | num zeros: 49063880.0 | params norm: 9732.894 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 5:17:02.031186 | finish at 2025-09-18 14:53:03 |
| [2025-09-18 09:52:42] iteration 41/ 59 | consumed samples: 41984 | elapsed time per iteration (ms): 1001076.8 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 7.041134E-06 | global batch size: 1024 | lm loss: 5.951908E+00 | loss scale: 1.0 | grad norm: 159719407616.000 | num zeros: 47581900.0 | params norm: 9732.887 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 5:00:19.381758 | finish at 2025-09-18 14:53:02 |
| [2025-09-18 10:09:23] iteration 42/ 59 | consumed samples: 43008 | elapsed time per iteration (ms): 1000910.9 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 6.700295E-06 | global batch size: 1024 | lm loss: 5.992692E+00 | loss scale: 1.0 | grad norm: 198826213376.000 | num zeros: 49029128.0 | params norm: 9732.880 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 4:43:35.484690 | finish at 2025-09-18 14:52:59 |
| [2025-09-18 10:26:04] iteration 43/ 59 | consumed samples: 44032 | elapsed time per iteration (ms): 1001109.6 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 6.359456E-06 | global batch size: 1024 | lm loss: 6.063922E+00 | loss scale: 1.0 | grad norm: 81292369920.000 | num zeros: 48821880.0 | params norm: 9732.874 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 4:26:57.753551 | finish at 2025-09-18 14:53:02 |
| [2025-09-18 10:42:45] iteration 44/ 59 | consumed samples: 45056 | elapsed time per iteration (ms): 1001102.0 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 6.018617E-06 | global batch size: 1024 | lm loss: 6.145180E+00 | loss scale: 1.0 | grad norm: 217426198528.000 | num zeros: 49321256.0 | params norm: 9732.868 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 4:10:16.529392 | finish at 2025-09-18 14:53:02 |
| [2025-09-18 10:59:26] iteration 45/ 59 | consumed samples: 46080 | elapsed time per iteration (ms): 1000832.6 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 5.677779E-06 | global batch size: 1024 | lm loss: 6.157466E+00 | loss scale: 1.0 | grad norm: 96730759168.000 | num zeros: 49428652.0 | params norm: 9732.861 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 3:53:31.656602 | finish at 2025-09-18 14:52:58 |
| [2025-09-18 11:16:07] iteration 46/ 59 | consumed samples: 47104 | elapsed time per iteration (ms): 1001193.6 | throughput per GPU (TFLOP/s/GPU): 284.6 | MFU 28.78% | learning rate: 5.336939E-06 | global batch size: 1024 | lm loss: 6.188946E+00 | loss scale: 1.0 | grad norm: 666173046784.000 | num zeros: 48664652.0 | params norm: 9732.856 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 3:36:55.516861 | finish at 2025-09-18 14:53:03 |
| [2025-09-18 11:32:48] iteration 47/ 59 | consumed samples: 48128 | elapsed time per iteration (ms): 1001008.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 4.996101E-06 | global batch size: 1024 | lm loss: 6.225647E+00 | loss scale: 1.0 | grad norm: 185815629824.000 | num zeros: 49034972.0 | params norm: 9732.851 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 3:20:12.100107 | finish at 2025-09-18 14:53:01 |
| [2025-09-18 11:49:30] iteration 48/ 59 | consumed samples: 49152 | elapsed time per iteration (ms): 1001083.1 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 4.655262E-06 | global batch size: 1024 | lm loss: 6.273437E+00 | loss scale: 1.0 | grad norm: 534045065216.000 | num zeros: 48900832.0 | params norm: 9732.846 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 3:03:31.914597 | finish at 2025-09-18 14:53:01 |
| [2025-09-18 12:06:10] iteration 49/ 59 | consumed samples: 50176 | elapsed time per iteration (ms): 1000901.7 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 4.314423E-06 | global batch size: 1024 | lm loss: 6.338528E+00 | loss scale: 1.0 | grad norm: 899446276096.000 | num zeros: 48965872.0 | params norm: 9732.841 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 2:46:49.016547 | finish at 2025-09-18 14:52:59 |
| [2025-09-18 12:22:51] iteration 50/ 59 | consumed samples: 51200 | elapsed time per iteration (ms): 1000783.5 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 3.973584E-06 | global batch size: 1024 | lm loss: 6.234092E+00 | loss scale: 1.0 | grad norm: 199767703552.000 | num zeros: 49578360.0 | params norm: 9732.838 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 2:30:07.051476 | finish at 2025-09-18 14:52:58 |
| [2025-09-18 12:39:32] iteration 51/ 59 | consumed samples: 52224 | elapsed time per iteration (ms): 1000899.7 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 3.632745E-06 | global batch size: 1024 | lm loss: 6.373925E+00 | loss scale: 1.0 | grad norm: 253297819648.000 | num zeros: 49478376.0 | params norm: 9732.834 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 2:13:27.197994 | finish at 2025-09-18 14:52:59 |
| [2025-09-18 12:56:13] iteration 52/ 59 | consumed samples: 53248 | elapsed time per iteration (ms): 1001001.0 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 3.291906E-06 | global batch size: 1024 | lm loss: 6.419476E+00 | loss scale: 1.0 | grad norm: 113803444224.000 | num zeros: 49280460.0 | params norm: 9732.830 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 1:56:47.007308 | finish at 2025-09-18 14:53:00 |
| [2025-09-18 13:12:54] iteration 53/ 59 | consumed samples: 54272 | elapsed time per iteration (ms): 1001045.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 2.951067E-06 | global batch size: 1024 | lm loss: 6.387888E+00 | loss scale: 1.0 | grad norm: 255267913728.000 | num zeros: 47073088.0 | params norm: 9732.827 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 1:40:06.271739 | finish at 2025-09-18 14:53:00 |
| [2025-09-18 13:29:35] iteration 54/ 59 | consumed samples: 55296 | elapsed time per iteration (ms): 1001091.3 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.78% | learning rate: 2.610229E-06 | global batch size: 1024 | lm loss: 6.394367E+00 | loss scale: 1.0 | grad norm: 131194019840.000 | num zeros: 49299976.0 | params norm: 9732.824 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 1:23:25.456586 | finish at 2025-09-18 14:53:01 |
| [2025-09-18 13:46:16] iteration 55/ 59 | consumed samples: 56320 | elapsed time per iteration (ms): 1000805.8 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 2.269390E-06 | global batch size: 1024 | lm loss: 6.428160E+00 | loss scale: 1.0 | grad norm: 526357069824.000 | num zeros: 49432300.0 | params norm: 9732.822 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 1:06:43.223324 | finish at 2025-09-18 14:52:59 |
| [2025-09-18 14:02:57] iteration 56/ 59 | consumed samples: 57344 | elapsed time per iteration (ms): 1000877.8 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.928551E-06 | global batch size: 1024 | lm loss: 6.425514E+00 | loss scale: 1.0 | grad norm: 80378552320.000 | num zeros: 49106088.0 | params norm: 9732.819 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 0:50:02.633369 | finish at 2025-09-18 14:53:00 |
| [2025-09-18 14:19:38] iteration 57/ 59 | consumed samples: 58368 | elapsed time per iteration (ms): 1000950.9 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.587712E-06 | global batch size: 1024 | lm loss: 6.411983E+00 | loss scale: 1.0 | grad norm: 601219596288.000 | num zeros: 49628192.0 | params norm: 9732.817 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 0:33:21.901740 | finish at 2025-09-18 14:53:00 |
| [2025-09-18 14:36:19] iteration 58/ 59 | consumed samples: 59392 | elapsed time per iteration (ms): 1000958.0 | throughput per GPU (TFLOP/s/GPU): 284.7 | MFU 28.79% | learning rate: 1.246873E-06 | global batch size: 1024 | lm loss: 6.455757E+00 | loss scale: 1.0 | grad norm: 586592092160.000 | num zeros: 48973652.0 | params norm: 9732.816 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 0:16:40.958041 | finish at 2025-09-18 14:53:00 |
| saving checkpoint at iteration 58 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 in torch format |
| successfully saved checkpoint from iteration 58 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 [ t 1/2, p 1/1 ] |
| (min, max) time across ranks (ms): |
| save-checkpoint ................................: (51718.88, 51718.90) |
| [2025-09-18 14:53:51] iteration 59/ 59 | consumed samples: 60416 | elapsed time per iteration (ms): 1000781.5 | throughput per GPU (TFLOP/s/GPU): 284.8 | MFU 28.79% | learning rate: 9.060344E-07 | global batch size: 1024 | lm loss: 6.509560E+00 | loss scale: 1.0 | grad norm: 141197361152.000 | num zeros: 49217292.0 | params norm: 9732.815 | number of skipped iterations: 0 | number of nan iterations: 0 | remaining time: 0:00:00 | finish at 2025-09-18 14:53:51 |
| [after training is done] datetime: 2025-09-18 14:53:51 |
| saving checkpoint at iteration 59 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 in torch format |
| successfully saved checkpoint from iteration 59 to /mnt/nanjingcephfs/project_wx-rec-alg-bdc-exp/bwzheng/yulan/hyw/pretrain-linear-moe-dev/megatron_lm_workspace/checkpoint/based-distill56l-dclm10b-s512-step394-mamba_hybrid-2.9b-112layers-q30-kv6-hybrid0.0625-pattern_A0-mheaddim64-mnumgroups6-mstatedim320-mexpand1-freeze_false-ep1-mp2-pp1-cp2-lr2e-5-minlr7e-7-bs1024-gpus8-seqlen32768 [ t 1/2, p 1/1 ] |
| [rank1]:[W918 14:54:44.101205233 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator()) |
| [rank3]:[W918 14:54:44.170021956 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator()) |
| [rank2]:[W918 14:54:44.807181589 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator()) |
| [rank0]:[W918 14:54:45.828822024 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator()) |
| [rank4]:[W918 14:54:45.679057092 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator()) |
| [rank6]:[W918 14:54:45.727992159 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator()) |
| [rank7]:[W918 14:54:45.747321323 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator()) |
| [rank5]:[W918 14:54:46.100354521 ProcessGroupNCCL.cpp:1476] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator()) |
|
|