--- language: - en license: apache-2.0 library_name: sentence-transformers tags: - sentence-transformers - feature-extraction - sentence-similarity - agents - rag - embeddings - llmops base_model: sentence-transformers/all-MiniLM-L6-v2 datasets: - hharsha/agentic-systems-showcase pipeline_tag: feature-extraction widget: - source_sentence: hybrid search and reranking for RAG sentences: - RetrievalLab shows advanced RAG with hybrid search and cross-encoder reranking. - A cooking recipe for pasta carbonara. - AgentFleet runs multi-agent task DAGs with cost governance. --- # agentic-systems-minilm Small **sentence embedding** model for semantic search over agentic / RAG / LLMOps project docs. > **Comparison note:** This is an **embeddings** model (MiniLM, 384-d), not a generative 7B demo > and not a Gradio chat Space. It is meant for retrieval and clustering on CPU. ## Model details | | | |---|---| | **Base** | [`sentence-transformers/all-MiniLM-L6-v2`](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) | | **Training** | 1 epoch CosineSimilarityLoss on CPU; pairs from [`hharsha/agentic-systems-showcase`](https://huggingface.co/datasets/hharsha/agentic-systems-showcase) plus short synthetic Q/A about AgentFleet, RetrievalLab, AgentOps Studio, Vibespace, Agent OS, CareerAgent, Control Tower, agentgrid | | **Intended use** | Semantic search / clustering of short texts about agentic systems, RAG pipelines, MCP tools, and related portfolio docs | | **Not intended for** | Open-domain chat, replacing large embedding models on broad corpora, medical/legal advice | | **License** | Apache-2.0 (base); training data MIT showcase | Light **domain adaptation** only — weights start from MiniLM-L6-v2. ## Usage ```python from sentence_transformers import SentenceTransformer model = SentenceTransformer("hharsha/agentic-systems-minilm") emb = model.encode(["hybrid search RAG with citations", "AgentFleet multi-agent ops"]) print(emb.shape) # (2, 384) ``` ## Related - Tag generator (text2text): [`hharsha/agentic-github-tagger`](https://huggingface.co/hharsha/agentic-github-tagger) - LoRA adapter (generative tiny): [`hharsha/agentic-rag-lora`](https://huggingface.co/hharsha/agentic-rag-lora) - Dataset: [`hharsha/agentic-systems-showcase`](https://huggingface.co/datasets/hharsha/agentic-systems-showcase) - Studio: [https://agentic-systems-studio.com](https://agentic-systems-studio.com) - GitHub: [https://github.com/hharsha98](https://github.com/hharsha98) ## How it was trained ```text Base: sentence-transformers/all-MiniLM-L6-v2 Loss: CosineSimilarityLoss Epochs: 1 (CPU) Batch size: 8 Data: name/summary pairs from the showcase dataset + synthetic agentic/RAG pairs ```