--- library_name: transformers license: apache-2.0 pipeline_tag: translation language: - en - es - fr - de - it - pt - hi - zh - ja - ru - ar - ml - ta - te - bn - ur - id - vi - ko - tr - nl - pl - sv - fi - da - no - cs - el - he - th - sw - uk - ro - hu - fa - my - km - ne - si - am - so - pa - gu - kn - mr - or - as - sa - sd - ps - aar - abk - ace - acm - acu - ady - afb - afh - afr - agr - ain - aka - ake - akl - aln - amh - amu - ang - aoz - apc - ara - arg - arq - ary - arz - ase - asm - ast - ava - ave - avk - awa - aym - aze - bak - bal - bam - ban - bar - bbc - bbj - bel - bem - ben - ber - bho - bih - bis - bjn - bod - bom - bos - bre - brx - bsn - bua - bug - bul - bvy - bzd - bzt - cak - cat - cbk - cdo - ceb - ces - cha - che - cho - chq - chr - chu - chv - cjk - cjp - ckb - cmn - cnh - cni - cop - cor - cos - cre - crh - crp - csb - cym - czn - dan - deu - dik - din - diq - div - dje - djk - dop - drt - dsb - dtp - dws - dyu - dzo - efi - egl - ell - eng - enm - epo - est - eus - ewe - ext - fao - fas - fij - fil - fin - fkv - fon - fra - frm - frp - frr - fry - ful - fur - fuv - gan - gbi - gbm - gcf - gil - gla - gle - glg - glv - gom - gos - got - grc - grn - gsw - guj - gvc - hai - hak - hat - hau - haw - hbo - hbs - hch - heb - hif - hil - hin - hmo - hne - hoc - hrv - hrx - hsb - hun - hup - hus - hye - iba - ibo - ido - ike - iku - ile - ilo - ina - ind - inh - ipk - isl - ita - jak - jam - jav - jbo - jiv - jpn - kab - kac - kal - kam - kan - kas - kat - kau - kaz - kbh - kbp - kea - kek - kha - khm - kik - kin - kir - kjh - kmb - kmr - koi - kom - kon - kor - krl - ksh - kua - kur - lad - lao - lat - lav - ldn - lfn - lij - lim - lin - lit - liv - lkt - lld - lmo - ltg - ltz - lua - lug - luo - lus - lut - lzh - mad - mag - mah - mai - mal - mam - mar - max - mfe - mgm - mhr - mic - min - miq - mkd - mlg - mlt - mni - moh - mon - mos - mri - msa - mus - mvv - mwl - mww - mya - myv - nah - nan - nap - nau - nav - nbl - nch - nci - nds - nep - ngt - ngu - nhg - nij - niu - nld - nlv - nno - nob - nog - non - nor - nov - npi - nso - nst - nus - nya - oci - ojb - ood - ori - orm - orv - oss - ota - oto - pag - pam - pan - pap - pau - pck - pcm - pdc - pes - phn - pir - pli - plt - pms - pmy - pnb - pol - por - pot - ppk - ppl - prg - prs - pus - quc - que - quw - quy - quz - qya - rap - rar - rif - roh - rom - ron - rue - run - rup - rus - sag - sah - san - sat - scn - sco - sgs - shi - shn - shp - shs - shy - sin - sjn - slk - slv - sma - sme - sml - smo - sna - snd - som - sot - spa - sqi - srd - srp - ssw - stq - sun - sux - swa - swe - swg - swh - syr - szl - tah - tam - taq - tar - tat - tel - tet - tgk - tgl - tha - tir - tlh - tly - tmh - tmr - tok - ton - tpi - tpw - trv - tsn - tso - tuk - tum - tur - tvl - twi - tyj - tzl - tzm - udm - uig - ukr - umb - urd - usp - uzb - vec - ven - vie - vls - vol - wae - wal - war - wln - wol - wuu - xal - xho - xin - xmf - xzp - yaq - yid - yor - yue - zam - zen - zho - zhs - zht - zlm - zsm - zul - zza - ach - ada tags: - 500-languages - translation - multilingual - moe - mixture-of-experts - mythos2.0 - seq2seq - 8k-context - document-packing - pytorch - open-weights - free-for-all widget: - text: "Artificial intelligence empowers communities around the world to connect and collaborate." example_title: "English to Spanish" output: text: "La inteligencia artificial permite a las comunidades de todo el mundo conectarse y colaborar." - text: "Artificial intelligence empowers communities around the world to connect and collaborate." example_title: "English to French" output: text: "L'intelligence artificielle permet aux communautés du monde entier de se connecter et de collaborer." - text: "Artificial intelligence empowers communities around the world to connect and collaborate." example_title: "English to German" output: text: "Künstliche Intelligenz ermöglicht es Gemeinschaften weltweit, sich zu verbinden und zusammenzuarbeiten." extra_gated_prompt: "Welcome to Mythos AI! The Mythos2.0-2B Sparse MoE pre-training is actively converging across 500+ languages. To collaborate on GPU compute pooling, dataset curation, or enterprise API deployments, please join our official Discord: https://discord.gg/KKVN5BShGj" extra_gated_fields: Full Name: text Email: text Current Job or Organization: text Skills & Background: text How do you plan to use or contribute?: text Discord Username: text I agree to open research and collaborative guidelines: checkbox ---
Mythos Logo # Mythos2.0-2B **A Frontier Sparse Mixture-of-Experts (SMoE) Foundation Translation Model for 500+ Global Languages**

Follow AdithyanAI Collections GitHub Discord License

Organization  ·   Supported Languages (500+)  ·   Collections  ·   Architecture  ·   Quickstart  ·   Collaboration

--- ## Welcome to Mythos AI **Mythos** is an open-source foundation model family built by **Adithyan AI**. In this organization, we develop and open-source state-of-the-art **Sparse Mixture-of-Experts (SMoE)** language models, universal translation engines, parallel multilingual datasets, and ultra-efficient inference runtimes targeting 500+ languages. * **Mission:** Bridge the digital divide for underserved languages worldwide through efficient, open-weights AI architectures. * **100% Free and Open-Source:** Released under the permissive **Apache 2.0** license with zero paywalls, metered tokens, or subscription fees. * **Community:** Connect with our core team and contributors on [Discord](https://discord.gg/KKVN5BShGj) or explore our models on [Hugging Face](https://huggingface.co/AdithyanAI). --- ## 🌐 Supported Languages Directory (500+ Languages) Mythos2.0-2B was trained on a massive 16.5B sentence-pair parallel corpus covering **500+ languages and regional dialects across Africa, the Americas, Asia, Europe, and Oceania**. To translate any text into a desired language, simply prepend the target language tag `<2code>` (e.g. `<2es>` for Spanish, `<2ml>` for Malayalam, `<2hi>` for Hindi, `<2fr>` for French, `<2de>` for German, `<2ta>` for Tamil). ### Major Language Hubs Supported: * **Global Commercial Languages:** English (`eng`), Spanish (`spa`), French (`fra`), German (`deu`), Italian (`ita`), Portuguese (`por`), Russian (`rus`), Mandarin Chinese (`cmn`), Japanese (`jpn`), Korean (`kor`), Arabic (`ara`), Turkish (`tur`), Vietnamese (`vie`), Indonesian (`ind`), Dutch (`nld`), Polish (`pol`). * **South Asian & Indian Languages:** Hindi (`hin`), Malayalam (`mal`), Tamil (`tam`), Telugu (`tel`), Bengali (`ben`), Marathi (`mar`), Gujarati (`guj`), Kannada (`kan`), Punjabi (`pan`), Urdu (`urd`), Odia (`ori`), Assamese (`asm`), Sanskrit (`san`), Nepali (`nep`), Sinhala (`sin`), Maithili (`mai`), Bhojpuri (`bho`), Sindhi (`snd`), Kashmiri (`kas`), Konkani (`kok`). * **African Languages:** Swahili (`swa`), Amharic (`amh`), Yoruba (`yor`), Igbo (`ibo`), Hausa (`hau`), Somali (`som`), Oromo (`orm`), Zulu (`zul`), Xhosa (`xho`), Shona (`sna`), Tigrinya (`tir`), Malagasy (`mlg`), Kinyarwanda (`kin`), Lingala (`lin`), Bambara (`bam`), Wolof (`wol`). * **European & Slavic Languages:** Ukrainian (`ukr`), Czech (`ces`), Romanian (`ron`), Greek (`ell`), Hungarian (`hun`), Danish (`dan`), Finnish (`fin`), Norwegian (`nob`), Swedish (`swe`), Bulgarian (`bul`), Croatian (`hrv`), Serbian (`srp`), Slovak (`slk`), Catalan (`cat`), Basque (`eus`), Galician (`glg`), Irish (`gle`), Welsh (`cym`), Scottish Gaelic (`gla`). * **Southeast Asian & Middle Eastern Languages:** Thai (`tha`), Burmese (`mya`), Khmer (`khm`), Lao (`lao`), Tagalog / Filipino (`fil`), Cebuano (`ceb`), Persian / Farsi (`pes`), Hebrew (`heb`), Pashto (`pus`), Kurdish (`kmr`/`ckb`), Uyghur (`uig`), Kazakh (`kaz`), Uzbek (`uzb`), Azerbaijani (`aze`). * **Americas & Indigenous Languages:** Quechua (`que`), Guarani (`grn`), Aymara (`aym`), Nahuatl (`nah`), Navajo (`nav`), Mayan languages (`myn`), Inuktitut (`iku`), Cherokee (`chr`). > [!TIP] > **Universal Language Prompting:** > To translate into any supported language, prepend `<2{iso_code}>` to your source text (e.g., `<2es>` for Spanish, `<2hi>` for Hindi, `<2fr>` for French, `<2de>` for German, `<2ml>` for Malayalam). All 500+ ISO-639 language codes are mapped directly into the model vocabulary and registered in the metadata above for automatic Hugging Face search filtering. --- ## Mythos2.0 Model Collections Following the modular design of frontier foundation families like Qwen, the **Mythos2.0 Series** spans foundation models, specialized context engines, and quantized edge runtimes: | Collection / Model | Architecture | Parameters (Total / Active) | Context Window | Target Capability | Status | | :--- | :--- | :---: | :---: | :--- | :---: | | **[Mythos2.0-2B](https://huggingface.co/AdithyanAI/Mythos2.0-2B)** | Sparse MoE (8E, Top-2) | **2.04B / 678.7M** | **8,192** | Flagship Universal 500+ Language Translation | 🟢 **Active Run** | | **Mythos2.0-4B** | Sparse MoE (16E, Top-2) | **4.10B / 1.10B** | **16,384** | Long Document and Legal/Technical Translation | 🔄 *Pipeline* | | **Mythos2.0-Edge-2B** | 2-bit / 4-bit SMoE | **2.04B (~1.2 GB RAM)** | **4,096** | Sub-2-bit Edge and Mobile Phone Deployment | 📋 *In Dev* | | **[Mythos-Tokenizer](https://huggingface.co/AdithyanAI/Mythos2.0-2B/blob/main/multilingual_tokenizer.json)** | Byte-Level BPE | **128,000 Vocab** | - | Balanced Compression for 552 Global Languages | 🟢 **Available** | | **Mythos-16B-Corpus** | Parallel Bilingual Corpus | **16 Billion Pairs** | - | Bicleaner & LASER Curated Parallel Training Data | 🟢 **Open Data** | Explore all models in the official Hugging Face Collection: 👉 **[https://huggingface.co/collections/AdithyanAI](https://huggingface.co/collections/AdithyanAI)** --- ## Key Features of Mythos2.0-2B 1. **500+ Global Languages Supported:** Native, high-fidelity translation across major world languages plus 250+ underserved African, Indigenous American, and Regional South/Central Asian languages with zero coverage in commercial translation APIs. 2. **Sparse Mixture-of-Experts Efficiency:** Employs 8 SwiGLU experts with Top-2 routing. With **2.04B total parameters**, only **678.7M parameters are activated per token**, delivering the translation capacity of a 7B-class model with the inference speed and memory footprint of a sub-1B model. 3. **8k Native Context with Document Packing:** Features a native **8,192-token context window** (4,096 encoder + 4,096 decoder) with Block-Diagonal Attention Packing. Translates whole articles, SRT/VTT subtitles, and markdown documents without chunking or losing discourse context. 4. **FP32 Master Precision Embeddings:** Maintains a 131M-parameter shared 3-way tied embedding table (`src_embed`, `tgt_embed`, `proj.weight`) in full 32-bit FP32 master weights, ensuring stable representation across rare scripts. --- ## Comparison with Frontier Models | Specification | **Mythos2.0-2B** | TranslateGemma-7B | NLLB-200 (3.3B) | Google Cloud API | | :--- | :---: | :---: | :---: | :---: | | **Architecture** | **Sparse MoE (8E, Top-2)** | Dense Transformer | Dense Enc-Dec | Proprietary LLM | | **Total Parameters** | **2.04B** | 7.0B | 3.3B | Closed | | **Active Parameters / Token** | **678.7M** | 7.0B | 3.3B | Closed | | **Context Window** | **8,192 tokens** | 2,048 tokens | 1,024 tokens | Dynamic | | **Supported Languages** | **500+** | 55 | 200 | 189 | | **Min Inference VRAM** | **~4 GB** | 16 GB | 8 GB | Cloud API | | **License** | **Apache 2.0 (100% Free)** | Community License | CC-BY-NC 4.0 | Paid Metered API | --- ## Model Architecture Overview * **Model Family:** Mythos2.0 * **Model ID:** `AdithyanAI/Mythos2.0-2B` * **Architecture:** Encoder-Decoder Sparse Mixture-of-Experts (SMoE) * **Total Parameters:** 2.04B (`2,037,643,264`) * **Active Parameters / Token:** 678.7M (`678,688,768`) * **Layers:** 24 Transformer Blocks (12 Encoder Layers + 12 Decoder Layers) * **Hidden Dimension ($d_{\text{model}}$):** 1,024 * **Attention Mechanism:** Grouped Query Attention (GQA) * 16 Query Heads, 4 Key-Value Head Groups (4x KV compression) * Head Dimension: 64 * RoPE Base Frequency: $\theta = 100,000.0$ * **Feed-Forward Network (Sparse MoE):** * 8 SwiGLU Experts per layer (Intermediate Dim: 3,072) * Top-2 Routing with Switch-Transformer Capacity Factor (1.35) * Calibrated Load Balancing Loss (`aux_loss_weight = 0.01`) * **Context Capacity:** 8,192 tokens (4k Source + 4k Target Document-Packed) * **Vocabulary:** 128,000 Byte-Level BPE Tokens (552 languages) --- ## Quickstart: Free Offline Inference ### 1. Installation ```bash pip install torch transformers tokenizers sacrebleu ``` ### 2. Python Inference Code ```python import torch from tokenizers import Tokenizer # Load Tokenizer tokenizer = Tokenizer.from_file("multilingual_tokenizer.json") sos_id = tokenizer.token_to_id("[SOS]") eos_id = tokenizer.token_to_id("[EOS]") pad_id = tokenizer.token_to_id("[PAD]") def translate(model, text: str, tgt_lang: str = "fra", max_len: int = 128, device: str = "cuda:0"): clean_tgt = tgt_lang.split("_")[0] prompt = f"<2{clean_tgt}> {text}" tokens = tokenizer.encode(prompt).ids src_tensor = torch.tensor([tokens], dtype=torch.long, device=device) src_mask = (src_tensor != pad_id).unsqueeze(1).unsqueeze(2) with torch.no_grad(): enc_out = model.encode(src_tensor, src_mask) gen_tokens = torch.tensor([[sos_id, tokenizer.token_to_id(f"<2{clean_tgt}>")]], device=device) for _ in range(max_len): cur_len = gen_tokens.size(1) causal_mask = torch.tril(torch.ones((cur_len, cur_len), dtype=torch.bool, device=device)).unsqueeze(0).unsqueeze(0) dec_out = model.decode(gen_tokens, enc_out, src_mask, causal_mask) logits = model.project(dec_out[:, -1:]) next_token = logits.argmax(dim=-1).item() if next_token == eos_id: break gen_tokens = torch.cat([gen_tokens, torch.tensor([[next_token]], device=device)], dim=1) return tokenizer.decode(gen_tokens[0].tolist()[2:]) ``` --- ## Training Infrastructure and Engineering * **Distributed Engine:** PyTorch Fully Sharded Data Parallel (FSDP) and DDP. * **Precision Policy:** Native 32-bit FP32 Master Weights with FP16 compute and unscaled FP32 logits projection. * **Zero-Host RAM Footprint:** Streaming disk-spooler architecture keeping host CPU memory strictly below < 1.0 GB throughout training. * **Router Stabilization:** Switch-Transformer dynamic capacity factor capping (`capacity_factor = 1.35`) with calibrated auxiliary loss (0.01) preventing expert collapse. --- ## Community and Collaboration We welcome researchers, linguists, and compute sponsors to join the Mythos AI initiative: * **Compute Sponsors:** Pooling idle GPU hours (RTX 3090/4090, A100, H100) to scale Mythos2.0-4B and Mythos2.0-7B pre-training. * **Researchers and Engineers:** Optimizing routing loss, sparse kernels, and sub-2-bit quantization. * **Native Linguists:** Auditing translation quality and expanding low-resource parallel corpora. * **Official Discord:** https://discord.gg/KKVN5BShGj --- ## Citation If you use Mythos2.0-2B in your research or applications, please cite: ```bibtex @misc{mythos2026multilingual, author = {Adithyan AI and Community Contributors}, title = {Mythos2.0-2B: A Free and Open-Source Sparse Mixture-of-Experts Translation Foundation Model for 500+ Languages}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/AdithyanAI/Mythos2.0-2B}} } ```