Text Generation
Safetensors
English
talkie
gptq
4-bit precision
quantized
instruction-tuned
vintage-language-model
chat
custom_code
Instructions to use dtestnyrr/talkie-1930-13b-it-gptq-int4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- vLLM
How to use dtestnyrr/talkie-1930-13b-it-gptq-int4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dtestnyrr/talkie-1930-13b-it-gptq-int4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dtestnyrr/talkie-1930-13b-it-gptq-int4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/dtestnyrr/talkie-1930-13b-it-gptq-int4
- SGLang
How to use dtestnyrr/talkie-1930-13b-it-gptq-int4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dtestnyrr/talkie-1930-13b-it-gptq-int4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dtestnyrr/talkie-1930-13b-it-gptq-int4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dtestnyrr/talkie-1930-13b-it-gptq-int4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dtestnyrr/talkie-1930-13b-it-gptq-int4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use dtestnyrr/talkie-1930-13b-it-gptq-int4 with Docker Model Runner:
docker model run hf.co/dtestnyrr/talkie-1930-13b-it-gptq-int4
| """GPTQModel adapter for the Talkie architecture. | |
| Importing this module registers TalkieQModel under model_type='talkie' in | |
| GPTQModel's MODEL_MAP, so `GPTQModel.load(...)` and `GPTQModel.from_quantized(...)` | |
| work without manual configuration. | |
| Auto-detect produces the same module_tree, so this is purely for the from_quantized | |
| path (which doesn't run auto-detect — module_tree must be a class attribute). | |
| """ | |
| from __future__ import annotations | |
| from gptqmodel.models.base import BaseQModel | |
| from gptqmodel.models.auto import MODEL_MAP, SUPPORTED_MODELS | |
| class TalkieQModel(BaseQModel): | |
| # talkie uses functional F.rms_norm with no learnable scale, so there's no | |
| # named pre-lm-head normalization module. Empty string disables that hook. | |
| pre_lm_head_norm_module = "" | |
| # Module tree maps GPTQModel's iteration onto our TalkieDecoderLayer: | |
| # model.layers.{i}.self_attn.{q_proj,k_proj,v_proj,o_proj} | |
| # model.layers.{i}.mlp.{gate_proj,up_proj,down_proj} | |
| # Suffix :0/:1 declares quantization grouping order — q/k/v share input | |
| # (the post-attn-rmsnorm hidden state), o has a different input (SDPA output). | |
| # Same for gate/up sharing input vs down. Mirrors LlamaQModel. | |
| module_tree = [ | |
| "model", | |
| "layers", | |
| "#", | |
| { | |
| "self_attn": ("q_proj:0", "k_proj:0", "v_proj:0", "o_proj:1"), | |
| "mlp": ("gate_proj:0", "up_proj:0", "down_proj:1"), | |
| }, | |
| ] | |
| # Register under model_type='talkie' so GPTQModel.load auto-routes to us. | |
| MODEL_MAP["talkie"] = TalkieQModel | |
| if "talkie" not in SUPPORTED_MODELS: | |
| SUPPORTED_MODELS.append("talkie") | |