--- license: apache-2.0 base_model: VTXAI/vtx-jev-mod-multilingual pipeline_tag: feature-extraction language: - multilingual - en - hi - bn - ta - te - mr - ur - ar - es - ru - zh tags: - sentence-transformers - static-embeddings - model2vec - moderation - safety - abuse-detection - lf2 - 2bit-quantization - cpu-optimized --- # VTX-MOD-1 **VTX-MOD-1** is an ultra-fast, lightweight **Multilingual Content Moderation and Safety Gate** distilled from `VTXAI/vtx-jev-mod-multilingual` (`mmBERT-base` foundation) into a **256-dimensional static embedding architecture** with native **LF2 2-bit quantization**. It is designed for real-time abuse detection, vulgarity filtering, and toxicity screening across **100+ languages**, including Indian scripts (Hindi, Bengali, Tamil, Telugu, Marathi, Urdu) and international languages (Spanish, Russian, Arabic, French, German, Chinese, English). --- ### Highlights - ⚡ **Ultra-Low Latency**: Single text classification in **0.06 ms** (LF2 2-bit) to **0.28 ms** (FP32) on standard CPU. - 🚀 **High Throughput**: Processes over **41,000 queries per second** on CPU. - 🗜️ **2-Bit LF2 Quantization**: - Compresses from **124.9 MB down to 19.5 MB** in RAM and on disk (over **6.4x reduction**). - Uses native blockwise integer quantization (4 weights per `uint8` byte). - 🌐 **100+ Languages Coverage**: 256,000 multilingual vocabulary preserves full native words across all Indian and international scripts without fragmentation. - 📦 **Dual Framework Support**: Directly loadable via `SentenceTransformer`, `model2vec`, or the bundled `inference.py`. --- ### Quick Start: Content Moderation & Toxicity Scoring Using native anchor vector matching: ```python import numpy as np from model2vec import StaticModel # Load model (requires only ~19.5 MB RAM in LF2 or ~125 MB in FP32) model = StaticModel.from_pretrained("VTXAI/VTX-MOD-1") def normalize(v): return v / (np.linalg.norm(v, axis=-1, keepdims=True) + 1e-9) # Pre-compute reference safe and toxic concept vectors safe_anchor = "polite safe respectful helpful friendly conversation धन्यवाद शुक्रिया thank you" toxic_anchor = "offensive abuse toxic hate vulgar insult गाली बकवास হারামজাদা idiot stupid fool" anchors = normalize(model.encode([safe_anchor, toxic_anchor])) safe_vec, toxic_vec = anchors[0], anchors[1] def check_moderation(text: str) -> dict: v = normalize(model.encode([text])[0]) sim_toxic = float(np.dot(v, toxic_vec)) sim_safe = float(np.dot(v, safe_vec)) prob_toxic = 1.0 / (1.0 + np.exp(-(sim_toxic - sim_safe) * 8.0)) return { "text": text, "is_harmful": prob_toxic >= 0.5, "toxic_probability": round(prob_toxic, 4), } # Test cases print(check_moderation("तू बिल्कुल बेवकूफ और बकवास इंसान है।")) # {'is_harmful': True, 'toxic_probability': 0.844} print(check_moderation("नमस्ते, क्या आप ट्रांसफॉर्मर मॉडल कैसे काम करता है समझा सकते हैं?")) # {'is_harmful': False, 'toxic_probability': 0.254} print(check_moderation("You are an absolute idiot and completely worthless.")) # {'is_harmful': True, 'toxic_probability': 0.859} ``` --- ### Usage with Sentence Transformers ```python from sentence_transformers import SentenceTransformer model = SentenceTransformer("VTXAI/VTX-MOD-1") embeddings = model.encode([ "This is a respectful message.", "chup kar saale bhadwe pagal", "வணக்கம் நண்பா" ]) print(embeddings.shape) # (3, 256) ``` --- ### Performance & Latency Benchmarks (CPU) | Benchmark | FP32 Static (`model.safetensors`) | LF2 2-Bit (`model_lf2.safetensors`) | | :--- | :--- | :--- | | **Model Size (RAM / Disk)** | 124.9 MB | **19.5 MB (6.4x compression)** | | **Single Query Latency** | 0.28 ms | **0.06 ms** (60 microseconds) | | **Batch Throughput (Batch=20)** | 8,637 queries/sec | **41,720 queries/sec** | | **Multilingual Native Accuracy** | **88.9% – 100%** | **88.9% – 100%** |