Text Generation
Transformers
Safetensors
GGUF
English
qwen3
dpo
preference
rlhf
manim
manim-voiceover
aos
code-generation
animation
merged
conversational
text-generation-inference
Instructions to use nabin2004/AOS-qwen3-8b-narrated-merged with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nabin2004/AOS-qwen3-8b-narrated-merged") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("nabin2004/AOS-qwen3-8b-narrated-merged") model = AutoModelForCausalLM.from_pretrained("nabin2004/AOS-qwen3-8b-narrated-merged", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nabin2004/AOS-qwen3-8b-narrated-merged with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M # Run inference directly in the terminal: llama cli -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M # Run inference directly in the terminal: llama cli -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Use Docker
docker model run hf.co/nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use nabin2004/AOS-qwen3-8b-narrated-merged with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nabin2004/AOS-qwen3-8b-narrated-merged" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nabin2004/AOS-qwen3-8b-narrated-merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
- SGLang
How to use nabin2004/AOS-qwen3-8b-narrated-merged with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nabin2004/AOS-qwen3-8b-narrated-merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nabin2004/AOS-qwen3-8b-narrated-merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nabin2004/AOS-qwen3-8b-narrated-merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nabin2004/AOS-qwen3-8b-narrated-merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Ollama:
ollama run hf.co/nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
- Unsloth Desktop
- Pi
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Docker Model Runner:
docker model run hf.co/nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
- Lemonade
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Run and chat with the model
lemonade run user.AOS-qwen3-8b-narrated-merged-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use nabin2004/AOS-qwen3-8b-narrated-merged with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use nabin2004/AOS-qwen3-8b-narrated-merged with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "nabin2004/AOS-qwen3-8b-narrated-merged:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download README.md from nabin2004/AOS-qwen3-8b-narrated-merged: direct link, hf CLI and curl.
- Browser
- Download file 5.17 kB
-
https://huggingface.co/nabin2004/AOS-qwen3-8b-narrated-merged/resolve/5c09e6e65f4d5f0dbafb204389e07babb5ce9c37/README.md
- Command line
-
hf download hf://nabin2004/AOS-qwen3-8b-narrated-merged@5c09e6e65f4d5f0dbafb204389e07babb5ce9c37/README.md
-
curl -L -o README.md https://huggingface.co/nabin2004/AOS-qwen3-8b-narrated-merged/resolve/5c09e6e65f4d5f0dbafb204389e07babb5ce9c37/README.md
5.17 kB
| license: apache-2.0 | |
| base_model: Qwen/Qwen3-8B | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| tags: | |
| - safetensors | |
| - dpo | |
| - preference | |
| - rlhf | |
| - manim | |
| - manim-voiceover | |
| - aos | |
| - code-generation | |
| - animation | |
| - merged | |
| # AOS Qwen3 8B Narrated (DPO Aligned - Merged Safetensors) | |
| Direct Preference Optimization (DPO) aligned full bf16 model for **Manim Community Edition** mathematical and educational animation synthesis with synchronized voiceover narration. | |
| **Model Repository:** [nabin2004/AOS-qwen3-8b-narrated-merged](https://huggingface.co/nabin2004/AOS-qwen3-8b-narrated-merged) | |
| --- | |
| ## Lineage & Provenance | |
| | Role | Artifact / Repository | Description | | |
| |---|---|---| | |
| | **Base LLM** | [`Qwen/Qwen3-8B`](https://huggingface.co/Qwen/Qwen3-8B) | Base causal foundation model | | |
| | **SFT Prior** | [`nabin2004/AOS-qwen3-8b-narrated-adapter`](https://huggingface.co/nabin2004/AOS-qwen3-8b-narrated-adapter) | Continued SFT on 400 synchronized educational voiceover trajectories | | |
| | **DPO Adapter** | [`nabin2004/AOS-qwen3-8b-narrated-dpo`](https://huggingface.co/nabin2004/AOS-qwen3-8b-narrated-dpo) | Direct Preference Optimization adapter ($\beta=0.1$) | | |
| | **Merged Weights** | [`nabin2004/AOS-qwen3-8b-narrated-merged`](https://huggingface.co/nabin2004/AOS-qwen3-8b-narrated-merged) | Full unquantized bfloat16 Safetensors weights (this repo) | | |
| | **GGUF / Ollama** | [`nabin2004/AOS-qwen3-8b-narrated-gguf`](https://huggingface.co/nabin2004/AOS-qwen3-8b-narrated-gguf) | Multi-quantized GGUF (`Q4_K_M`, `Q8_0`) for Ollama & llama.cpp | | |
| --- | |
| ## Alignment Objective | |
| The model was aligned with Direct Preference Optimization (DPO) to strongly prefer generating voiceover-synchronized educational animations: | |
| - **Chosen**: Clean `VoiceoverScene` scripts with speech services (`AOSSpeechService` / `GTTSService`), animation duration tracking (`run_time=tracker.duration`), millisecond-accurate `<bookmark mark='...'/>` tags, and natural phonetic spoken narration. | |
| - **Rejected**: Silent, un-narrated standard `Scene` code. | |
| --- | |
| ## Canonical Code Pattern | |
| ```python | |
| from manim import * | |
| from manim_voiceover import VoiceoverScene | |
| from manim_voiceover.services.gtts import GTTSService | |
| class SigmoidExplanation(VoiceoverScene): | |
| def construct(self): | |
| # Configure speech service | |
| self.set_speech_service(GTTSService()) | |
| title = Title("The Sigmoid Activation Function") | |
| ax = Axes(x_range=[-6, 6, 2], y_range=[-0.2, 1.2, 0.5]) | |
| curve = ax.plot(lambda x: 1 / (1 + np.exp(-x)), color=BLUE) | |
| dot = Dot(ax.c2p(0, 0.5), color=RED) | |
| with self.voiceover( | |
| text="Let's visualize the sigmoid function. <bookmark mark='AXES'/> We begin by setting up our coordinate system, <bookmark mark='CURVE'/> plotting the characteristic S-shaped curve, <bookmark mark='DOT'/> and marking the midpoint inflection at zero, point five." | |
| ) as tracker: | |
| self.play(Write(title)) | |
| self.wait_until_bookmark("AXES") | |
| self.play(Create(ax)) | |
| self.wait_until_bookmark("CURVE") | |
| self.play(Create(curve)) | |
| self.wait_until_bookmark("DOT") | |
| self.play(FadeIn(dot), run_time=tracker.duration) | |
| self.wait(1) | |
| ``` | |
| --- | |
| ## Quickstart Usage | |
| ### Hugging Face Transformers | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "nabin2004/AOS-qwen3-8b-narrated-merged" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| torch_dtype=torch.bfloat16, | |
| device_map="auto", | |
| trust_remote_code=True, | |
| ) | |
| prompt = "Create a narrated Manim animation explaining the Fourier Transform with voiceover bookmarks." | |
| messages = [ | |
| {{"role": "system", "content": "You are an expert mathematical animation assistant specializing in Manim Community Edition and voiceover narration with manim-voiceover. You write complete, self-contained, fully executable Python scripts inheriting from VoiceoverScene."}}, | |
| {{"role": "user", "content": prompt}} | |
| ] | |
| text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) | |
| inputs = tokenizer([text], return_tensors="pt").to(model.device) | |
| outputs = model.generate( | |
| **inputs, | |
| max_new_tokens=2048, | |
| temperature=0.2, | |
| top_p=0.95, | |
| pad_token_id=tokenizer.eos_token_id, | |
| ) | |
| print(tokenizer.decode(outputs[0][len(inputs.input_ids[0]):], skip_special_tokens=True)) | |
| ``` | |
| ### High-Throughput Cloud Serving with vLLM | |
| Serve as a high-performance OpenAI-compatible endpoint: | |
| ```bash | |
| vllm serve nabin2004/AOS-qwen3-8b-narrated-merged \ | |
| --port 8000 \ | |
| --max-model-len 8192 \ | |
| --trust-remote-code | |
| ``` | |
| --- | |
| ## Citation & Acknowledgments | |
| Part of the **AOS (Agentic Orchestration System)** project for multi-agent educational video synthesis. | |
| - Base Model: Alibaba Cloud Qwen Team (`Qwen/Qwen3-8B`) | |
| - Animation Engine: Manim Community Edition & `manim-voiceover` | |