--- library_name: sentence-transformers pipeline_tag: sentence-similarity tags: - sentence-transformers - feature-extraction - sentence-similarity - transformers - echo - linear-complexity - recurrent-hybrid datasets: - sentence-transformers/all-nli - mteb/sts-b metrics: - spearman - pearson language: - en --- # Echo-DSRN Embedding Model (Echo-DSRN-v0.1.3-Embed-Exp) This is a high-performance sentence embedding model based on the recurrent-hybrid **Echo-DSRN** architecture. It scales **linearly** ($O(N)$) with sequence length, offering extreme efficiency and sub-millisecond latency on both CPU and GPU. ## 🚀 Model Details - **Developer:** ethicalabs - **Base Model:** [ethicalabs/Echo-DSRN-114M-v0.1.2](https://huggingface.co/ethicalabs/Echo-DSRN-114M-v0.1.2) - **Architecture:** [Echo-DSRN](https://huggingface.co/ethicalabs/Echo-DSRN-114M-v0.1.2) Recurrent-Hybrid - **Parameters:** 98.26M (98,264,064) - **Embedding Dimension:** 2048 - **Max Sequence Length:** 2048 ## 📊 Evaluation Results (MTEB STS) Spearman Rank Correlation scores on Semantic Textual Similarity (STS) benchmark tasks: | Benchmark Task | Echo-DSRN (Ours) | | :--- | :---: | | **STS12** | `0.6667` | | **STS13** | `0.7692` | | **STS14** | `0.7683` | | **STS15** | `0.8227` | | **STS16** | `0.7460` | | **STSBenchmark** | `0.7293` | | **SICK-R** | `0.7876` | | **Average STS** | **`0.7557`** | ## ⚡ Efficiency and Systems Scaling Profile Inference performance (latency and peak VRAM allocation) on GPU and CPU configurations across different sequence lengths: ### GPU Latency & VRAM Benchmark | Sequence Length | Echo-DSRN Latency (GPU) | Echo-DSRN VRAM (GPU) | | :---: | :---: | :---: | | 128 | `15.93 ms` | `516.69 MB` | | 256 | `17.56 ms` | `548.44 MB` | | 512 | `32.14 ms` | `604.95 MB` | | 1024 | `71.30 ms` | `710.96 MB` | | 2048 | `155.26 ms` | `932.99 MB` | | 4096 | N/A (OOR) | N/A (OOR) | ### CPU Latency Benchmark | Sequence Length | Echo-DSRN Latency (CPU) | | :---: | :---: | | 128 | `48.50 ms` | | 256 | `84.94 ms` | | 512 | `160.93 ms` | | 1024 | `328.93 ms` | | 2048 | `727.57 ms` | | 4096 | N/A (OOR) | *Note: 'N/A (OOR)' indicates sequence length exceeds model's maximum position embedding range.* ## 🏗️ Architecture Details | Property | Value | | :--- | :--- | | Layers | 8 | | Hidden Dim | 512 | | Vocab Size | 32017 | | Attention Heads | 4 | ## 📊 Parameter Breakdown | Component | Parameters | % of Total | | :--- | :--- | :--- | | **Total** | **98.26M (98,264,064)** | **100%** | | Embeddings | 16.39M | 16.68% | | DSRN Recurrent Blocks | 81.87M | 83.32% | | Norms & Biases | 512 | 0.00% | ## 💻 Usage You can load and use this model directly via `sentence-transformers`: ```python from sentence_transformers import SentenceTransformer # Load model with auto-mapping enabled model = SentenceTransformer("ethicalabs/Echo-DSRN-v0.1.3-Embed-Exp", trust_remote_code=True) # Encode text to get 2048-dimensional embeddings sentences = ["The recurrent slow state contains the aligned sequence representations.", "Echo-DSRN has linear complexity."] embeddings = model.encode(sentences) print(embeddings.shape) # (2, 2048) ``` ## 🛠️ Training Procedure The model was trained in three sequential phases: 1. **Contrastive Pre-training**: Representation space alignment using natural language inference datasets. 2. **Fine-grained Similarity Tuning**: Fine-tuning using semantic textual similarity benchmarks to calibrate similarity scores. 3. **Multi-Task Generalization Tuning**: Training on NLI retrieval, Banking77 intent classification, and STS semantic similarity simultaneously with dynamic early stopping to prevent representation collapse. --- *This model card was automatically generated by `scripts/generate_model_card.py`.*