File size: 3,332 Bytes
3ddb1a3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
---
language:
- en
license: mit
library_name: tokenizers
tags:
- sentence-similarity
- feature-extraction
- embeddings
- rag
- quantized
- 2-bit
- lf2
- matryoshka
- ultra-lightweight
- code-search
- retrieval
- vortexa
pipeline_tag: feature-extraction
---

<div align="center">

# 🚀 vtx-embed-1M-lf2 (`nano-2bit`)

**The world's most memory-efficient native 2-bit static embedding model powering [vortexa](https://github.com/OEvortex/vortexa).**
Native 2-Bit LF2 integer quantization · **0.39 MB RAM** · 1.05M Parameters · Fused Numba dequant & mean-pooling · Sub-millisecond CPU latency

[![HuggingFace](https://img.shields.io/badge/🤗%20HuggingFace-VTXAI%2Fvtx-embed-1M-blue)](https://huggingface.co/VTXAI/vtx-embed-1M)
[![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](https://opensource.org/licenses/MIT)
[![Python 3.8+](https://img.shields.io/badge/Python-3.8%2B-blue)](https://python.org)

</div>

---

## ⚡ What is LF2?

**LF2** is an ultra-compact, integer-native 2-bit quantization format for embedding matrices:
- **Zero FP32 Parameter Tables**: Parameters are stored entirely as packed 2-bit levels (4 weights per `uint8` byte) and double-quantized `uint8` scales and minimums.
- **Extreme Compression**: Memory drops from 0.57 MB (LF4 4-bit) down to **0.39 MB** (2-bit), retaining **95.40% mean token cosine similarity** to full precision.
- **Fused Dequantization & Pooling**: Evaluated on-the-fly using Numba JIT kernels directly into registers / L1 cache without allocating full FP32 token tables in RAM.

---

## 📄 Model Details

| Property | Value |
| :--- | :--- |
| **Model Name / Tier** | **vtx-embed-1M-lf2** (`"nano-2bit"`) |
| **Total Parameters** | **1.05M** |
| **Quantization Format** | `lf2` (Native 2-bit integer block quantization) |
| **In-RAM Memory** | **0.39 MB** |
| **On-Disk Size** | **0.39 MB** |
| **Embedding Dimension** | 64 |
| **Vocabulary Size** | 16,384 |
| **Block Size** | 16 |
| **Mean Cosine Similarity vs Orig** | **0.9540** |
| **License** | MIT |

---

## 💻 Quickstart Usage

### Standalone Inference with `lf2_native.py`

This repository includes [`lf2_native.py`](lf2_native.py) directly for zero-dependency inference (only requires `numpy`, `safetensors`, and `tokenizers`; `numba` optional for maximum speed):

```python
from huggingface_hub import snapshot_download
import sys

# 1. Download model repository
model_path = snapshot_download(repo_id="VTXAI/vtx-embed-1M-lf2")
sys.path.append(model_path)

from lf2_native import VortexEmbedLF2

# 2. Load model directly from local directory
model = VortexEmbedLF2.from_pretrained(model_path)
print(f"Model In-RAM size: {model.model_size_mb:.2f} MB")

# 3. Encode sentences
texts = [
    "What is the capital of India?",
    "Explain gravity and general relativity",
]
embeddings = model.encode(texts)
print("Embeddings shape:", embeddings.shape)  # (2, 64)

# 4. Semantic similarity
sim = embeddings[0] @ embeddings[1]
print("Cosine Similarity:", sim)
```

---

## 📜 Citation

```bibtex
@misc{vtx-embed-1m-lf2,
  title  = {vtx-embed-1M-lf2: Native 2-Bit Embeddings for Ultra-Low Footprint Semantic Search},
  author = {VTXAI},
  year   = {2026},
  url    = {https://huggingface.co/VTXAI/vtx-embed-1M-lf2}
}
```

---

## 📄 License

MIT License — free for commercial and research use.