--- language: - en license: mit library_name: tokenizers tags: - sentence-similarity - feature-extraction - embeddings - rag - quantized - 2-bit - lf2 - matryoshka - ultra-lightweight - code-search - retrieval - vortexa pipeline_tag: feature-extraction ---
# ๐Ÿš€ vtx-embed-1M-lf2 (`nano-2bit`) **The world's most memory-efficient native 2-bit static embedding model powering [vortexa](https://github.com/OEvortex/vortexa).** Native 2-Bit LF2 integer quantization ยท **0.39 MB RAM** ยท 1.05M Parameters ยท Fused Numba dequant & mean-pooling ยท Sub-millisecond CPU latency [![HuggingFace](https://img.shields.io/badge/๐Ÿค—%20HuggingFace-VTXAI%2Fvtx-embed-1M-blue)](https://huggingface.co/VTXAI/vtx-embed-1M) [![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](https://opensource.org/licenses/MIT) [![Python 3.8+](https://img.shields.io/badge/Python-3.8%2B-blue)](https://python.org)
--- ## โšก What is LF2? **LF2** is an ultra-compact, integer-native 2-bit quantization format for embedding matrices: - **Zero FP32 Parameter Tables**: Parameters are stored entirely as packed 2-bit levels (4 weights per `uint8` byte) and double-quantized `uint8` scales and minimums. - **Extreme Compression**: Memory drops from 0.57 MB (LF4 4-bit) down to **0.39 MB** (2-bit), retaining **95.40% mean token cosine similarity** to full precision. - **Fused Dequantization & Pooling**: Evaluated on-the-fly using Numba JIT kernels directly into registers / L1 cache without allocating full FP32 token tables in RAM. --- ## ๐Ÿ“„ Model Details | Property | Value | | :--- | :--- | | **Model Name / Tier** | **vtx-embed-1M-lf2** (`"nano-2bit"`) | | **Total Parameters** | **1.05M** | | **Quantization Format** | `lf2` (Native 2-bit integer block quantization) | | **In-RAM Memory** | **0.39 MB** | | **On-Disk Size** | **0.39 MB** | | **Embedding Dimension** | 64 | | **Vocabulary Size** | 16,384 | | **Block Size** | 16 | | **Mean Cosine Similarity vs Orig** | **0.9540** | | **License** | MIT | --- ## ๐Ÿ’ป Quickstart Usage ### Standalone Inference with `lf2_native.py` This repository includes [`lf2_native.py`](lf2_native.py) directly for zero-dependency inference (only requires `numpy`, `safetensors`, and `tokenizers`; `numba` optional for maximum speed): ```python from huggingface_hub import snapshot_download import sys # 1. Download model repository model_path = snapshot_download(repo_id="VTXAI/vtx-embed-1M-lf2") sys.path.append(model_path) from lf2_native import VortexEmbedLF2 # 2. Load model directly from local directory model = VortexEmbedLF2.from_pretrained(model_path) print(f"Model In-RAM size: {model.model_size_mb:.2f} MB") # 3. Encode sentences texts = [ "What is the capital of India?", "Explain gravity and general relativity", ] embeddings = model.encode(texts) print("Embeddings shape:", embeddings.shape) # (2, 64) # 4. Semantic similarity sim = embeddings[0] @ embeddings[1] print("Cosine Similarity:", sim) ``` --- ## ๐Ÿ“œ Citation ```bibtex @misc{vtx-embed-1m-lf2, title = {vtx-embed-1M-lf2: Native 2-Bit Embeddings for Ultra-Low Footprint Semantic Search}, author = {VTXAI}, year = {2026}, url = {https://huggingface.co/VTXAI/vtx-embed-1M-lf2} } ``` --- ## ๐Ÿ“„ License MIT License โ€” free for commercial and research use.