vtx-embed-1M-lf2 / README.md
Abhaykoul's picture
Initial release of native LF2 2-bit quantized embedding model
3ddb1a3 verified
|
Raw History Blame Contribute Delete
3.33 kB
metadata
language:
  - en
license: mit
library_name: tokenizers
tags:
  - sentence-similarity
  - feature-extraction
  - embeddings
  - rag
  - quantized
  - 2-bit
  - lf2
  - matryoshka
  - ultra-lightweight
  - code-search
  - retrieval
  - vortexa
pipeline_tag: feature-extraction

🚀 vtx-embed-1M-lf2 (nano-2bit)

The world's most memory-efficient native 2-bit static embedding model powering vortexa. Native 2-Bit LF2 integer quantization · 0.39 MB RAM · 1.05M Parameters · Fused Numba dequant & mean-pooling · Sub-millisecond CPU latency

HuggingFace License: MIT Python 3.8+


⚡ What is LF2?

LF2 is an ultra-compact, integer-native 2-bit quantization format for embedding matrices:

  • Zero FP32 Parameter Tables: Parameters are stored entirely as packed 2-bit levels (4 weights per uint8 byte) and double-quantized uint8 scales and minimums.
  • Extreme Compression: Memory drops from 0.57 MB (LF4 4-bit) down to 0.39 MB (2-bit), retaining 95.40% mean token cosine similarity to full precision.
  • Fused Dequantization & Pooling: Evaluated on-the-fly using Numba JIT kernels directly into registers / L1 cache without allocating full FP32 token tables in RAM.

📄 Model Details

Property Value
Model Name / Tier vtx-embed-1M-lf2 ("nano-2bit")
Total Parameters 1.05M
Quantization Format lf2 (Native 2-bit integer block quantization)
In-RAM Memory 0.39 MB
On-Disk Size 0.39 MB
Embedding Dimension 64
Vocabulary Size 16,384
Block Size 16
Mean Cosine Similarity vs Orig 0.9540
License MIT

💻 Quickstart Usage

Standalone Inference with lf2_native.py

This repository includes lf2_native.py directly for zero-dependency inference (only requires numpy, safetensors, and tokenizers; numba optional for maximum speed):

from huggingface_hub import snapshot_download
import sys

# 1. Download model repository
model_path = snapshot_download(repo_id="VTXAI/vtx-embed-1M-lf2")
sys.path.append(model_path)

from lf2_native import VortexEmbedLF2

# 2. Load model directly from local directory
model = VortexEmbedLF2.from_pretrained(model_path)
print(f"Model In-RAM size: {model.model_size_mb:.2f} MB")

# 3. Encode sentences
texts = [
    "What is the capital of India?",
    "Explain gravity and general relativity",
]
embeddings = model.encode(texts)
print("Embeddings shape:", embeddings.shape)  # (2, 64)

# 4. Semantic similarity
sim = embeddings[0] @ embeddings[1]
print("Cosine Similarity:", sim)

📜 Citation

@misc{vtx-embed-1m-lf2,
  title  = {vtx-embed-1M-lf2: Native 2-Bit Embeddings for Ultra-Low Footprint Semantic Search},
  author = {VTXAI},
  year   = {2026},
  url    = {https://huggingface.co/VTXAI/vtx-embed-1M-lf2}
}

📄 License

MIT License — free for commercial and research use.