How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "N-Bot-Int/OpenElla-NovelWriter-8B-merged" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "N-Bot-Int/OpenElla-NovelWriter-8B-merged",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "N-Bot-Int/OpenElla-NovelWriter-8B-merged" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "N-Bot-Int/OpenElla-NovelWriter-8B-merged",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Quick Links

Uploaded finetuned model

  • Developed by: N-Bot-Int
  • License: agp-3
  • Finetuned from model : p-e-w/Llama-3.1-8B-Instruct-heretic

This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.

AN EXPERIMENTAL AI MODEL MADE USING LLAMA 3.1 WITH THE SAME DATASET AS OTHER NEW MODEL RELEASED THIS APRIL(RPG MIXED V1-V2-V3), THE MODEL IS CURRENTLY UNDER EXPERIMENTATION!

Downloads last month
37
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for N-Bot-Int/OpenElla-NovelWriter-8B-merged

Merges
2 models
Quantizations
2 models

Collection including N-Bot-Int/OpenElla-NovelWriter-8B-merged