NeoMME logo

NeoMME-260M predecay checkpoint

Hugging Face Hugging Face arXiv

Model summary

This repository contains the NeoMME-260M checkpoint at step 450,000 of a 500,000-step pretraining run. As the model name suggests, this checkpoint was saved just before the learning rate decay phase. It is intended as a starting point for continued pretraining with NeoMMEForMaskedLM.

ItemValue
Training step450,000 of 500,000
Intended useContinued pretraining
Model classNeoMMEForMaskedLM
Training objectivePredict masked text tokens from visible text and image patches

training_state.pt contains the original native NeoMME optimizer and training state. It is provided for research and custom continuation workflows, but it is not directly compatible with Hugging Face Trainer or standard AdamW. The file uses NeoMME's custom NorMuon and MasterAdamW state formats, native parameter grouping, and FP32 master weights.

Continue pretraining

The original run used the following optimizer settings:

Parameter groupOptimizerPeak learning rateWeight decay
Two-dimensional weight matricesNorMuon0.0120.01
Embedding tablesAdamW0.00130.01
One-dimensional parametersAdamW0.00130

The original learning-rate schedule used:

  • A 300-step linear warmup.
  • A stable phase through step 450,000.
  • A 50,000-step linear decay from the peak learning rate to 1% of the peak learning rate.

For a Transformers continuation with new optimizers, initialize them at the peak learning rates and apply the 50,000-step linear decay.

License

Model weights are released under the Apache 2.0 license.

Citation

@misc{lac2026neommesingletowermultimodalnativemultilingual,
      title={NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference},
      author={Aurélien Lac and Tony Wu},
      year={2026},
      eprint={2609.01657},
      archivePrefix={arXiv},
      primaryClass={cs.IR},
      url={https://arxiv.org/abs/2609.01657},
}
Downloads last month
128
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including Hcompany/NeoMME-260M-Pretrain-predecay-s450000

Paper for Hcompany/NeoMME-260M-Pretrain-predecay-s450000

Article mentioning Hcompany/NeoMME-260M-Pretrain-predecay-s450000