|
Download README.md from VTXAI/vtx-embed-1M-lf2: direct link, hf CLI and curl.
- Browser
- Download file 3.33 kB
-
https://huggingface.co/VTXAI/vtx-embed-1M-lf2/resolve/main/README.md
- Command line
-
hf download hf://VTXAI/vtx-embed-1M-lf2/README.md
-
curl -L -o README.md https://huggingface.co/VTXAI/vtx-embed-1M-lf2/resolve/main/README.md
3.33 kB
metadata
language:
- en
license: mit
library_name: tokenizers
tags:
- sentence-similarity
- feature-extraction
- embeddings
- rag
- quantized
- 2-bit
- lf2
- matryoshka
- ultra-lightweight
- code-search
- retrieval
- vortexa
pipeline_tag: feature-extraction
🚀 vtx-embed-1M-lf2 (nano-2bit)
The world's most memory-efficient native 2-bit static embedding model powering vortexa. Native 2-Bit LF2 integer quantization · 0.39 MB RAM · 1.05M Parameters · Fused Numba dequant & mean-pooling · Sub-millisecond CPU latency
⚡ What is LF2?
LF2 is an ultra-compact, integer-native 2-bit quantization format for embedding matrices:
- Zero FP32 Parameter Tables: Parameters are stored entirely as packed 2-bit levels (4 weights per
uint8byte) and double-quantizeduint8scales and minimums. - Extreme Compression: Memory drops from 0.57 MB (LF4 4-bit) down to 0.39 MB (2-bit), retaining 95.40% mean token cosine similarity to full precision.
- Fused Dequantization & Pooling: Evaluated on-the-fly using Numba JIT kernels directly into registers / L1 cache without allocating full FP32 token tables in RAM.
📄 Model Details
| Property | Value |
|---|---|
| Model Name / Tier | vtx-embed-1M-lf2 ("nano-2bit") |
| Total Parameters | 1.05M |
| Quantization Format | lf2 (Native 2-bit integer block quantization) |
| In-RAM Memory | 0.39 MB |
| On-Disk Size | 0.39 MB |
| Embedding Dimension | 64 |
| Vocabulary Size | 16,384 |
| Block Size | 16 |
| Mean Cosine Similarity vs Orig | 0.9540 |
| License | MIT |
💻 Quickstart Usage
Standalone Inference with lf2_native.py
This repository includes lf2_native.py directly for zero-dependency inference (only requires numpy, safetensors, and tokenizers; numba optional for maximum speed):
from huggingface_hub import snapshot_download
import sys
# 1. Download model repository
model_path = snapshot_download(repo_id="VTXAI/vtx-embed-1M-lf2")
sys.path.append(model_path)
from lf2_native import VortexEmbedLF2
# 2. Load model directly from local directory
model = VortexEmbedLF2.from_pretrained(model_path)
print(f"Model In-RAM size: {model.model_size_mb:.2f} MB")
# 3. Encode sentences
texts = [
"What is the capital of India?",
"Explain gravity and general relativity",
]
embeddings = model.encode(texts)
print("Embeddings shape:", embeddings.shape) # (2, 64)
# 4. Semantic similarity
sim = embeddings[0] @ embeddings[1]
print("Cosine Similarity:", sim)
📜 Citation
@misc{vtx-embed-1m-lf2,
title = {vtx-embed-1M-lf2: Native 2-Bit Embeddings for Ultra-Low Footprint Semantic Search},
author = {VTXAI},
year = {2026},
url = {https://huggingface.co/VTXAI/vtx-embed-1M-lf2}
}
📄 License
MIT License — free for commercial and research use.