Instructions to use nabin2004/AOS-qwen3-8b-narrated-merged with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nabin2004/AOS-qwen3-8b-narrated-merged") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("nabin2004/AOS-qwen3-8b-narrated-merged") model = AutoModelForCausalLM.from_pretrained("nabin2004/AOS-qwen3-8b-narrated-merged", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nabin2004/AOS-qwen3-8b-narrated-merged with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M # Run inference directly in the terminal: llama cli -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M # Run inference directly in the terminal: llama cli -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Use Docker
docker model run hf.co/nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use nabin2004/AOS-qwen3-8b-narrated-merged with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nabin2004/AOS-qwen3-8b-narrated-merged" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nabin2004/AOS-qwen3-8b-narrated-merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
- SGLang
How to use nabin2004/AOS-qwen3-8b-narrated-merged with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nabin2004/AOS-qwen3-8b-narrated-merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nabin2004/AOS-qwen3-8b-narrated-merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nabin2004/AOS-qwen3-8b-narrated-merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nabin2004/AOS-qwen3-8b-narrated-merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Ollama:
ollama run hf.co/nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
- Unsloth Desktop
- Pi
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Docker Model Runner:
docker model run hf.co/nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
- Lemonade
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Run and chat with the model
lemonade run user.AOS-qwen3-8b-narrated-merged-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use nabin2004/AOS-qwen3-8b-narrated-merged with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
AOS Qwen3 8B Narrated (DPO Aligned - Merged Safetensors)
Direct Preference Optimization (DPO) aligned full bf16 model for Manim Community Edition mathematical and educational animation synthesis with synchronized voiceover narration.
Model Repository: nabin2004/AOS-qwen3-8b-narrated-merged
Lineage & Provenance
| Role | Artifact / Repository | Description |
|---|---|---|
| Base LLM | Qwen/Qwen3-8B |
Base causal foundation model |
| SFT Prior | nabin2004/AOS-qwen3-8b-narrated-adapter |
Continued SFT on 400 synchronized educational voiceover trajectories |
| DPO Adapter | nabin2004/AOS-qwen3-8b-narrated-dpo |
Direct Preference Optimization adapter ($\beta=0.1$) |
| Merged Weights | nabin2004/AOS-qwen3-8b-narrated-merged |
Full unquantized bfloat16 Safetensors weights (this repo) |
| GGUF / Ollama | nabin2004/AOS-qwen3-8b-narrated-gguf |
Multi-quantized GGUF (Q4_K_M, Q8_0) for Ollama & llama.cpp |
Alignment Objective
The model was aligned with Direct Preference Optimization (DPO) to strongly prefer generating voiceover-synchronized educational animations:
- Chosen: Clean
VoiceoverScenescripts with speech services (AOSSpeechService/GTTSService), animation duration tracking (run_time=tracker.duration), millisecond-accurate<bookmark mark='...'/>tags, and natural phonetic spoken narration. - Rejected: Silent, un-narrated standard
Scenecode.
Canonical Code Pattern
from manim import *
from manim_voiceover import VoiceoverScene
from manim_voiceover.services.gtts import GTTSService
class SigmoidExplanation(VoiceoverScene):
def construct(self):
# Configure speech service
self.set_speech_service(GTTSService())
title = Title("The Sigmoid Activation Function")
ax = Axes(x_range=[-6, 6, 2], y_range=[-0.2, 1.2, 0.5])
curve = ax.plot(lambda x: 1 / (1 + np.exp(-x)), color=BLUE)
dot = Dot(ax.c2p(0, 0.5), color=RED)
with self.voiceover(
text="Let's visualize the sigmoid function. <bookmark mark='AXES'/> We begin by setting up our coordinate system, <bookmark mark='CURVE'/> plotting the characteristic S-shaped curve, <bookmark mark='DOT'/> and marking the midpoint inflection at zero, point five."
) as tracker:
self.play(Write(title))
self.wait_until_bookmark("AXES")
self.play(Create(ax))
self.wait_until_bookmark("CURVE")
self.play(Create(curve))
self.wait_until_bookmark("DOT")
self.play(FadeIn(dot), run_time=tracker.duration)
self.wait(1)
Quickstart Usage
Hugging Face Transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "nabin2004/AOS-qwen3-8b-narrated-merged"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
prompt = "Create a narrated Manim animation explaining the Fourier Transform with voiceover bookmarks."
messages = [
{{"role": "system", "content": "You are an expert mathematical animation assistant specializing in Manim Community Edition and voiceover narration with manim-voiceover. You write complete, self-contained, fully executable Python scripts inheriting from VoiceoverScene."}},
{{"role": "user", "content": prompt}}
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=2048,
temperature=0.2,
top_p=0.95,
pad_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(outputs[0][len(inputs.input_ids[0]):], skip_special_tokens=True))
High-Throughput Cloud Serving with vLLM
Serve as a high-performance OpenAI-compatible endpoint:
vllm serve nabin2004/AOS-qwen3-8b-narrated-merged \
--port 8000 \
--max-model-len 8192 \
--trust-remote-code
Citation & Acknowledgments
Part of the AOS (Agentic Orchestration System) project for multi-agent educational video synthesis.
- Base Model: Alibaba Cloud Qwen Team (
Qwen/Qwen3-8B) - Animation Engine: Manim Community Edition &
manim-voiceover
- Downloads last month
- 493