File size: 2,727 Bytes
906f4fa
 
 
 
 
 
 
 
 
 
 
 
cba64e6
906f4fa
 
 
 
 
 
 
 
 
 
 
 
 
 
cba64e6
 
 
 
906f4fa
 
 
 
 
 
cba64e6
906f4fa
cba64e6
 
906f4fa
cba64e6
906f4fa
 
 
 
 
 
 
 
 
 
 
cba64e6
 
 
 
 
 
 
 
906f4fa
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
---
language:
- en
license: apache-2.0
library_name: sentence-transformers
tags:
- sentence-transformers
- feature-extraction
- sentence-similarity
- agents
- rag
- embeddings
- llmops
base_model: sentence-transformers/all-MiniLM-L6-v2
datasets:
- hharsha/agentic-systems-showcase
pipeline_tag: feature-extraction
widget:
- source_sentence: hybrid search and reranking for RAG
  sentences:
  - RetrievalLab shows advanced RAG with hybrid search and cross-encoder reranking.
  - A cooking recipe for pasta carbonara.
  - AgentFleet runs multi-agent task DAGs with cost governance.
---

# agentic-systems-minilm

Small **sentence embedding** model for semantic search over agentic / RAG / LLMOps project docs.

> **Comparison note:** This is an **embeddings** model (MiniLM, 384-d), not a generative 7B demo
> and not a Gradio chat Space. It is meant for retrieval and clustering on CPU.

## Model details

| | |
|---|---|
| **Base** | [`sentence-transformers/all-MiniLM-L6-v2`](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) |
| **Training** | 1 epoch CosineSimilarityLoss on CPU; pairs from [`hharsha/agentic-systems-showcase`](https://huggingface.co/datasets/hharsha/agentic-systems-showcase) plus short synthetic Q/A about AgentFleet, RetrievalLab, AgentOps Studio, Vibespace, Agent OS, CareerAgent, Control Tower, agentgrid |
| **Intended use** | Semantic search / clustering of short texts about agentic systems, RAG pipelines, MCP tools, and related portfolio docs |
| **Not intended for** | Open-domain chat, replacing large embedding models on broad corpora, medical/legal advice |
| **License** | Apache-2.0 (base); training data MIT showcase |

Light **domain adaptation** only — weights start from MiniLM-L6-v2.

## Usage

```python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("hharsha/agentic-systems-minilm")
emb = model.encode(["hybrid search RAG with citations", "AgentFleet multi-agent ops"])
print(emb.shape)  # (2, 384)
```

## Related

- Tag generator (text2text): [`hharsha/agentic-github-tagger`](https://huggingface.co/hharsha/agentic-github-tagger)
- LoRA adapter (generative tiny): [`hharsha/agentic-rag-lora`](https://huggingface.co/hharsha/agentic-rag-lora)
- Dataset: [`hharsha/agentic-systems-showcase`](https://huggingface.co/datasets/hharsha/agentic-systems-showcase)
- Studio: [https://agentic-systems-studio.com](https://agentic-systems-studio.com)
- GitHub: [https://github.com/hharsha98](https://github.com/hharsha98)

## How it was trained

```text
Base: sentence-transformers/all-MiniLM-L6-v2
Loss: CosineSimilarityLoss
Epochs: 1 (CPU)
Batch size: 8
Data: name/summary pairs from the showcase dataset + synthetic agentic/RAG pairs
```