Instructions to use TensorFold/Qwen3.8-27B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TensorFold/Qwen3.8-27B-MLX-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("TensorFold/Qwen3.8-27B-MLX-4bit") config = load_config("TensorFold/Qwen3.8-27B-MLX-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use TensorFold/Qwen3.8-27B-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Qwen3.8-27B-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TensorFold/Qwen3.8-27B-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use TensorFold/Qwen3.8-27B-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Qwen3.8-27B-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TensorFold/Qwen3.8-27B-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TensorFold/Qwen3.8-27B-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Qwen3.8-27B-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TensorFold/Qwen3.8-27B-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
awesome, some questions
First I love your work, I recognize the name and I go ther ealong with a few others.
I have the M3U 256 gig too which is near what you have but architecturally the same more or less, so it helps me get a preview of what to expect + your configurations and instructions with oMLX.
The questions I have are the following, hoping you may answer:
Is there a real difference between q4 and q6 or q8/9? As I understand it, q4 is worse on smaller models, potentially more so with MLX (though that may be outdated due to MLX advancements). Q4 is better on larger models, perhaps 70b+? We can run models easily at Q8, but is Q6 the real sweet spot for larger ram + bandwidth users? I've also read that on paper a larger quant should be better but can be worse even.
Do you plan on establishing at least oMLX benchmarks? It is a lot to ask, so I'll be running on my own but the biggest data you provide that is most helpful is the highlevel such as decode speed - its rarely shared but THANK YOU for the effort here alone.
Will you make an oq4/6/8 version - you've uplaoded these, ty!
Would you recommend setting KV cache to Q8 in oMLX?
Are there any configurations for MTP, DSPARK, etc.? Sorry, I'm taking a quick rbeak from work so haven't been able to review in fine details. I noticed you just upload MLX q6 and q8 - cool!
Please always include great instructions and configs for both CLI and in-app settings of the oMLx app - great work.
Thanks, really appreciate it.
Yes, there is a real difference between Q4, Q6, and Q8, but it varies a lot by model.
On 256 GB systems, I generally like Q6 as the sweet spot. Q4 is great when memory or speed matters, while Q8 is useful when you want to minimize quantization loss.
Yes, I want to add more oMLX benchmarks.
Decode speed, memory use, context, and exact settings are probably the most useful numbers to share.
Yep, I’ve started uploading oQ4/Q6/Q8 where it makes sense.
With 256 GB, I’d generally start with Q8 KV cache. I’d only go lower if long context or concurrency makes memory an issue.
I’m also looking at MTP/DSpark configs. oMLX is changing quickly, so I’d rather test them properly before recommending specific settings.
And agreed, I’ll keep including both CLI commands and oMLX app settings with uploads.
for me the bottleneck is prefill speed and memory, not cache - the cache itself is much smaller than what it needs for prefill (not that it makes much sense to me tbh)
TQ slows down the model a bit, so whether you need it is up to you. If you run a single agent that mostly runs serialized requests or chats then there's not much benefit in enabling TQ. If you run many different workloads with numerous varying prefixes then TQ will help and 8bit will have absolutely negligible impact but save 50% of memory (but only for cache, not prefill)
oQ4/oQ4e works quite well for dense model like 27B, but avoid it for MoE (35B-A3B etc.) - there is degrades a lot and the lowest quant I'd consider is oQ6
MTP/DFlash is up to how it benchmarks on your machine, in my experience MTP is simply better, faster, with lower memory overhead. DFlash is sometimes faster for short prompts but you lose MTP - I'd rather keep MTP for everything