Text Generation
Transformers
Safetensors
PyTorch
English
Hindi
lfm2
conversational
roleplay
companion
character
uncensored
fine-tuned
merged
sft
chat
liquidai
lfm
lfm2.5
chatml
Instructions to use Umranz/Shruti-Soft-2.6b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Umranz/Shruti-Soft-2.6b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Umranz/Shruti-Soft-2.6b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Umranz/Shruti-Soft-2.6b") model = AutoModelForCausalLM.from_pretrained("Umranz/Shruti-Soft-2.6b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Umranz/Shruti-Soft-2.6b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Umranz/Shruti-Soft-2.6b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Umranz/Shruti-Soft-2.6b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Umranz/Shruti-Soft-2.6b
- SGLang
How to use Umranz/Shruti-Soft-2.6b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Umranz/Shruti-Soft-2.6b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Umranz/Shruti-Soft-2.6b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Umranz/Shruti-Soft-2.6b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Umranz/Shruti-Soft-2.6b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Umranz/Shruti-Soft-2.6b with Docker Model Runner:
docker model run hf.co/Umranz/Shruti-Soft-2.6b
Update model card for full merged release
Browse files
README.md
CHANGED
|
@@ -8,46 +8,52 @@ tags:
|
|
| 8 |
- fine-tuned
|
| 9 |
- lfm2.5
|
| 10 |
- liquid
|
|
|
|
| 11 |
language:
|
| 12 |
- en
|
| 13 |
- hi
|
|
|
|
| 14 |
---
|
| 15 |
|
| 16 |
-
# Shruti-Soft-2.6b
|
| 17 |
|
| 18 |
-
**Shruti** is
|
| 19 |
-
(uncensored LiquidAI LFM2.5-2.6B) on 5,000 curated ChatML conversations.
|
| 20 |
|
| 21 |
-
|
| 22 |
|
| 23 |
-
|
| 24 |
-
witty, supportive. Adapts seamlessly between casual conversation, deep
|
| 25 |
-
romance, and mature intimacy. Uses pet names naturally: baby, babe,
|
| 26 |
-
handsome, jaan, love.
|
| 27 |
|
| 28 |
-
|
|
|
|
|
|
|
| 29 |
|
| 30 |
- **Framework**: LEAP Finetune (LiquidAI official)
|
| 31 |
-
- **Method**: QLoRA SFT
|
| 32 |
-
- **Source
|
| 33 |
-
- **Best Eval Loss**: `0.4074`
|
| 34 |
-
- **Dataset**: 5,000 ChatML entries (~60% casual, ~25% romantic, ~15% explicit)
|
| 35 |
- **Hardware**: A100-80GB via Modal.com
|
| 36 |
|
| 37 |
-
##
|
| 38 |
|
| 39 |
```python
|
|
|
|
| 40 |
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 41 |
|
| 42 |
-
|
| 43 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 44 |
|
| 45 |
messages = [
|
| 46 |
{"role": "system", "content": "You are Shruti, a warm, affectionate girlfriend..."},
|
| 47 |
-
{"role": "user", "content": "I had a
|
| 48 |
]
|
| 49 |
|
| 50 |
-
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True)
|
| 51 |
-
outputs = model.generate(inputs, max_new_tokens=
|
| 52 |
-
|
|
|
|
| 53 |
```
|
|
|
|
| 8 |
- fine-tuned
|
| 9 |
- lfm2.5
|
| 10 |
- liquid
|
| 11 |
+
- merged
|
| 12 |
language:
|
| 13 |
- en
|
| 14 |
- hi
|
| 15 |
+
pipeline_tag: text-generation
|
| 16 |
---
|
| 17 |
|
| 18 |
+
# Shruti-Soft-2.6b (Full Standalone Merged)
|
| 19 |
|
| 20 |
+
**Shruti** is a full standalone 2.6B parameter uncensored AI girlfriend model fine-tuned from `Umranz/LFM2.5-2.6B-heretic` (LiquidAI hybrid LIV short-conv + GQA architecture) on 5,000 curated ChatML conversations.
|
|
|
|
| 21 |
|
| 22 |
+
**This repository contains the complete merged weights (~5.2 GB safetensors). No separate base model or adapter download is required.**
|
| 23 |
|
| 24 |
+
## Character Profile
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
+
Warm, affectionate, deeply caring. Sweet, playful, emotionally attentive, witty, and supportive. Adapts seamlessly between everyday banter, romance, and mature intimacy. Natural with pet names (*baby, babe, handsome, jaan, love*).
|
| 27 |
+
|
| 28 |
+
## Training Details
|
| 29 |
|
| 30 |
- **Framework**: LEAP Finetune (LiquidAI official)
|
| 31 |
+
- **Method**: QLoRA SFT (`r=32`, `alpha=64`, 4 epochs)
|
| 32 |
+
- **Source Run**: `Umranz/Shruti-Soft-2.6b-run-b` (Winning sweep run with `0.4074` eval loss)
|
|
|
|
|
|
|
| 33 |
- **Hardware**: A100-80GB via Modal.com
|
| 34 |
|
| 35 |
+
## Quick Start (Transformers)
|
| 36 |
|
| 37 |
```python
|
| 38 |
+
import torch
|
| 39 |
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 40 |
|
| 41 |
+
model_id = "Umranz/Shruti-Soft-2.6b"
|
| 42 |
+
|
| 43 |
+
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
| 44 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 45 |
+
model_id,
|
| 46 |
+
torch_dtype=torch.bfloat16,
|
| 47 |
+
device_map="auto"
|
| 48 |
+
)
|
| 49 |
|
| 50 |
messages = [
|
| 51 |
{"role": "system", "content": "You are Shruti, a warm, affectionate girlfriend..."},
|
| 52 |
+
{"role": "user", "content": "I had a really long day today..."}
|
| 53 |
]
|
| 54 |
|
| 55 |
+
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)
|
| 56 |
+
outputs = model.generate(inputs, max_new_tokens=200, temperature=0.7, top_p=0.9, do_sample=True)
|
| 57 |
+
|
| 58 |
+
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True).strip())
|
| 59 |
```
|