--- library_name: transformers pipeline_tag: text-generation tags: - olmo3 - custom-code - siamese-norm - depth-attention - sliding-window-attention --- # OLMo 3 with SiameseNorm and DepthAttention This repository is one of five standalone OLMo 3 checkpoints trained with SiameseNorm and DepthAttention. All model, configuration, tokenizer, and custom modeling files are stored directly at the repository root. | Repository suffix | Training phase | Context length | |---|---|---:| | `stage1` | Pretraining | 8,192 | | `stage2` | Mid-training | 8,192 | | `stage3` | Long-context training | 65,536 | | `stage4-think` | Think SFT | 65,536 | | `stage4-instruct` | Instruct SFT | 65,536 | Stage 3 and Stage 4 apply YaRN only to Full-attention layers. Sliding-window attention layers retain the original RoPE and a 4,096-token window. ## Loading Set `repo_id` to the desired repository from the table above: ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo_id = "ArchSpace-Collection/OLMo3-SiameseNorm-DepthAttention-" tokenizer = AutoTokenizer.from_pretrained( repo_id, trust_remote_code=True, fix_mistral_regex=False, ) model = AutoModelForCausalLM.from_pretrained( repo_id, trust_remote_code=True, dtype=torch.bfloat16, attn_implementation="sdpa", ) ``` `fix_mistral_regex=False` preserves the tokenizer behavior used during training. SDPA is the recommended and release-validated BF16 inference backend. ## Architecture - 16 transformer layers - hidden size 2,048 - intermediate size 8,192 - 16 query heads and 16 key/value heads - 128-dimensional attention heads - 3:1 sliding-window/full-attention pattern - 4,096-token sliding window - reordered RMSNorm, SiameseNorm, and DepthAttention