How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "Retreatcost/Limn-Alpha-12B" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "Retreatcost/Limn-Alpha-12B",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "Retreatcost/Limn-Alpha-12B" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "Retreatcost/Limn-Alpha-12B",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Quick Links

Limn-Alpha-12B

Limn-Alpha-12B

A generalist finetune, more vivid language and different thinking patterns.

I've transplanted lm_head from Gryphe/Gemma-4-12B-StyleTune to uncensored base llmfan46/gemma-4-12B-it-uncensored-heretic and did an rsLoRA finetuning on top.

After a brief testing it seems that both audio and video capabilites have been preserved, in some cases outputs are even more descriptive.

Inference Tips

  1. Temperature: 1.0
  2. TOP_P: 0.95
  3. TOP_K: 0 (disable)
  4. MIN_P: 0.025
  5. Template Format: Gemma4

Both thinking and non-thinking works.

These settings are practically the same as official recommendations, but i found to like MinP better than TopK.

Training details

Spoiler warning Trained on same dataset as Evertide-RX-12B.

Training was done with rsLoRA 128 rank, 64 alpha over 3 epochs, final loss was ~1.22.

During training o_proj was omitted to preserve uncensoring from heretic base model.

Special Thanks

  • Gryphe Padar: for his cool StyleTune idea and model (and all other cool tunes he done).
  • LLMfan46: for his never-ending flow of high-quality herectic models.
  • Team mradermacher: for awesome quants in GGUF format

Future Plans

This is an "Alpha" tune, that turned out to be quite good in my opinion. Probably going to do an FFT with better sample composition and extended dataset.

Downloads last month
53
Safetensors
Model size
13B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Retreatcost/Limn-Alpha-12B

Finetuned
(3)
this model
Quantizations
2 models

Collection including Retreatcost/Limn-Alpha-12B