Instructions to use Ttimofeyka/Tissint-14B-v1.1-128k-RP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Ttimofeyka/Tissint-14B-v1.1-128k-RP with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Ttimofeyka/Tissint-14B-v1.1-128k-RP") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Ttimofeyka/Tissint-14B-v1.1-128k-RP") model = AutoModelForCausalLM.from_pretrained("Ttimofeyka/Tissint-14B-v1.1-128k-RP", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Ttimofeyka/Tissint-14B-v1.1-128k-RP with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Ttimofeyka/Tissint-14B-v1.1-128k-RP" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ttimofeyka/Tissint-14B-v1.1-128k-RP", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Ttimofeyka/Tissint-14B-v1.1-128k-RP
- SGLang
How to use Ttimofeyka/Tissint-14B-v1.1-128k-RP with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Ttimofeyka/Tissint-14B-v1.1-128k-RP" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ttimofeyka/Tissint-14B-v1.1-128k-RP", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Ttimofeyka/Tissint-14B-v1.1-128k-RP" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ttimofeyka/Tissint-14B-v1.1-128k-RP", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use Ttimofeyka/Tissint-14B-v1.1-128k-RP with Docker Model Runner:
docker model run hf.co/Ttimofeyka/Tissint-14B-v1.1-128k-RP
Tissint-14B-v1.1-128k-RP
The model is based on SuperNova-Medius (as the current best 14B model) with a 128k context with an emphasis on creativity, including NSFW and multi-turn conversations.
According to my tests, this finetune is much more stable with different samplers than the original model. Censorship and refusals have been reduced.
The model started to follow the system prompt better, and the responses in ChatML format with bad samplers stopped reaching 800+ tokens for no reason.
V1.1
I have increased the amount of dataset and added instructions mixed with NSFW RP, which in theory will improve the quality of the model.
Chat Template - ChatML
Samplers
Balance
Temp : 0.8 - 1.15
Min P : 0.1
Repetition Penalty : 1.02
DRY 0.8, 1.75, 2, 2048 (change to 4096 or more if needed)
Creativity
Temp : 1.15 - 1.5
Top P : 0.9
Repetition Penalty : 1.03
DRY 0.82, 1.75, 2, 2048 (change to 4096 or more if needed)
- Downloads last month
- 13
Model tree for Ttimofeyka/Tissint-14B-v1.1-128k-RP
Base model
Qwen/Qwen2.5-14B