--- base_model: dhanesh-hf/Jarvis-Titan-V14-MoE-Merged tags: - deepseek-moe - titans-neural-memory - pallas-tpu - ultra-long-context - sliding-window-attention - tri-brid - jarvis-titan - distillation - reasoning - agentic library_name: transformers pipeline_tag: text-generation license: other license_name: jtrl-v1.0 license_link: LICENSE --- # J.A.R.V.I.S. TITAN 14.8B MoE — MILESTONE M3 ULTRA-LONG ADAPTER Official Milestone M3 (Phase 3) weights for **J.A.R.V.I.S. Titan 14.8B DeepSeekMoE + Tri-Brid Memory Architecture**, distilled from full dense attention to Sliding Window Attention ($W=2048$) on Google Cloud TPU v5e-8. ## Distillation & Training Specifications - **Base Model**: `dhanesh-hf/Jarvis-Titan-V14-MoE-Merged` (14.75B MoE, 100% Frozen) - **Adapter Initialization**: `dhanesh-hf/Jarvis-Titan-M2-TriBrid-Adapter` (Phase 2 Tri-Brid) - **Dataset Source**: `dhanesh-hf/jarvis-v10-rft-dataset` (100% Real Non-NIAH Peer-Reviewed Papers & Code) - **Sliding Window Size**: $W = 2048$ tokens (capping KV cache at ~115 MB) - **Strategic Layers**: [3, 7, 11, 15, 19, 23, 27] (7 Memory Bridges) - **Tier 2 Salient Reservoir**: $R=1024$ slots ($H_Q=28, H_{KV}=4$) - **Tier 3 Titans Neural Memory**: $d=512$, Google Pallas TPU VMEM SRAM kernel - **Triple-Gated Adaptive Fusion**: MAG-3 ($g_{\text{local}}, g_{\text{res}}, g_{\text{mem}}$) - **Distillation Loss**: $(1 - 0.5) \mathcal{L}_{\text{CE}} + 0.5 T^2 \mathcal{D}_{\text{KL}}$ ($T=2.0$) - **Total Adapter Parameters**: 90,044,458 (90.04M) - **Total Tokens Trained**: 20,012,495