oktayd's picture
Add nested Ubuntu and Windows installations inside all four hardware quickstarts
b54a798 verified
|
Raw History Blame Contribute Delete
19.7 kB

๐Ÿš€ Start on your device

Choose your hardware, then expand Ubuntu/Linux or Windows/PowerShell inside it. You only need one edition and one quantization. Installation examples are not new benchmark results.

๐Ÿ–ฅ๏ธ H200 / large server โ€” full BF16 with Transformers

For a machine with enough GPU memory. Validated on Ubuntu / NVIDIA H200 with Transformers 5.16.1. BF16 weights alone are about 70 GB, plus runtime memory. A Windows client does not provide the remote server's VRAM; this package does not fit a single RTX 5090 or an 8-GB laptop GPU.

๐Ÿง Ubuntu / Linux โ€” installation

Install Python 3.11, virtual-environment support and a CUDA-enabled PyTorch build from the official PyTorch selector that matches your NVIDIA driver. On a fresh Ubuntu machine:

sudo apt-get update
sudo apt-get install -y python3-venv python3-pip
mkdir -p qwen-v2-server
cd qwen-v2-server
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
# Install the CUDA-enabled PyTorch command from the official selector here.
python -m pip install transformers==5.16.1 accelerate pillow
python -c "import torch; print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0))"

If the GPU check fails, fix the driver/PyTorch installation before downloading the model. Save the shared Python example below as qwen_chat.py, then run python qwen_chat.py. The first run downloads the BF16 weights.

๐ŸชŸ Windows / PowerShell โ€” native installation

Native Windows โ€” WSL2 is not required. Install Python 3.11 and a compatible NVIDIA driver, then use PowerShell 7. This native installation route is separate from the measured H200 validation, which was performed on Ubuntu:

New-Item -ItemType Directory -Path .\qwen-v2-server -Force | Out-Null
Set-Location .\qwen-v2-server
py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade pip
# Run the Windows CUDA PyTorch install command from the official selector,
# using .\.venv\Scripts\python.exe -m pip instead of pip.
.\.venv\Scripts\python.exe -m pip install transformers==5.16.1 accelerate pillow
.\.venv\Scripts\python.exe -c "import torch; print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0))"

Save the shared example below as qwen_chat.py, then execute .\.venv\Scripts\python.exe .\qwen_chat.py. Use only Windows GPU/driver combinations supported by PyTorch and with sufficient memory.

If the H200 belongs to a remote Ubuntu server, Windows is just the client. Connect using Windows OpenSSH (replace the example and use your actual port/key), then follow the Ubuntu instructions in that remote shell:

ssh user@YOUR_SERVER
๐Ÿง Optional: Ubuntu under WSL2

Optional alternative, not required for native Windows. Follow Microsoft's WSL installation guide. From an administrator PowerShell, only if you want this alternative:

wsl --install -d Ubuntu

Restart if requested and complete Ubuntu's first-run account setup. Then open Ubuntu:

wsl -d Ubuntu

Inside Ubuntu, follow this device's Ubuntu/Linux instructions and create a separate Linux Python environment. Keep the Windows NVIDIA driver current; use NVIDIA's CUDA-on-WSL guidance rather than installing a Linux display driver inside WSL. Verify GPU access before loading weights.

Use either the Windows or WSL model server during the test, not both on the same port. WSL does not add VRAM. Reuse model downloads where practical, but do not reuse Windows Python environments inside Linux; Ollama stores/imports may create another copy. Native Windows remains the default path above.

Shared Python example โ€” use after either installation:

import torch
from transformers import AutoProcessor, Qwen3_5MoeForConditionalGeneration

repo = "oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-FT"
processor = AutoProcessor.from_pretrained(repo)
model = Qwen3_5MoeForConditionalGeneration.from_pretrained(
    repo, dtype=torch.bfloat16, device_map="auto"
)
messages = [{"role": "user", "content": [
    {"type": "text", "text": "Explain gravity briefly."}
]}]
inputs = processor.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_dict=True, return_tensors="pt", enable_thinking=False
).to(model.device)
with torch.inference_mode():
    output = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(processor.batch_decode(
    output[:, inputs.input_ids.shape[1]:], skip_special_tokens=True
)[0])

For images, include an image content item supported by AutoProcessor. Large images and long contexts increase memory use.

๐ŸŽฎ RTX 5090 ยท 32 GB โ€” Q4_K_M with llama.cpp

Use the GGUF / Llama edition. Q4_K_M is about 20.22 GiB, plus the 0.84-GiB projector and runtime memory. RTX 5090 / Q4_K_M was benchmarked with llama.cpp commit 427291b5b34cd914a31b3fd3b61a68f6184f4b9f on Ubuntu. The Windows steps below are installation guidance, not a separate Windows 5090 result.

๐Ÿง Ubuntu / Linux โ€” installation & launch

Install Python, the HF CLI and a CUDA-enabled llama.cpp build. Put llama-server on PATH. Choose a driver/toolkit/build supporting the RTX 5090; a CPU-only binary is not sufficient for GPU offload.

mkdir -p qwen-v2-q4
cd qwen-v2-q4
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade huggingface_hub
llama-server --list-devices
repo="oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Llama"
hf download "$repo" Qwen3.6-35B-v2-Q4_K_M.gguf mmproj-Qwen3.6-35B-v2-F16.gguf SHA256SUMS LICENSE --local-dir .
sha256sum --check --ignore-missing SHA256SUMS
llama-server -m Qwen3.6-35B-v2-Q4_K_M.gguf   --mmproj mmproj-Qwen3.6-35B-v2-F16.gguf   -ngl 99 -c 4096 -np 1 --host 127.0.0.1 --port 8080   --chat-template-kwargs '{"enable_thinking":false}'

Stop if checksums fail or the device listing does not show the intended NVIDIA GPU. In another terminal:

curl http://127.0.0.1:8080/v1/chat/completions   -H "Content-Type: application/json"   -d '{"messages":[{"role":"user","content":"Explain gravity briefly."}],"max_tokens":256,"temperature":0,"stream":false}'
๐ŸชŸ Windows / PowerShell โ€” native installation & launch

Install Python 3.11 and download the matching Windows CUDA package and any required CUDA runtime DLL package from the official llama.cpp releases. Extract the full package, preserving its DLLs; add that folder to PATH for this terminal. Do not substitute a CPU-only build. These commands assume llama-server.exe is on PATH. Use PowerShell 7 for native JSON argument handling.

New-Item -ItemType Directory -Path .\qwen-v2-q4 -Force | Out-Null
Set-Location .\qwen-v2-q4
py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade huggingface_hub
llama-server.exe --list-devices
$qwenRepo = "oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Llama"
.\.venv\Scripts\hf.exe download $qwenRepo Qwen3.6-35B-v2-Q4_K_M.gguf mmproj-Qwen3.6-35B-v2-F16.gguf SHA256SUMS LICENSE --local-dir .
if ($LASTEXITCODE -ne 0) { throw "Download failed" }
$qwenFiles = @("Qwen3.6-35B-v2-Q4_K_M.gguf", "mmproj-Qwen3.6-35B-v2-F16.gguf")
$qwenChecks = @{}
Get-Content .\SHA256SUMS | ForEach-Object {
    if ($_ -match '^([0-9a-fA-F]{64})  (.+)$') { $qwenChecks[$matches[2]] = $matches[1] }
}
foreach ($qwenFile in $qwenFiles) {
    if (-not $qwenChecks.ContainsKey($qwenFile)) { throw "Missing checksum: $qwenFile" }
    if ((Get-FileHash -LiteralPath $qwenFile -Algorithm SHA256).Hash -ne $qwenChecks[$qwenFile]) { throw "Checksum mismatch: $qwenFile" }
}
llama-server.exe -m Qwen3.6-35B-v2-Q4_K_M.gguf --mmproj mmproj-Qwen3.6-35B-v2-F16.gguf -ngl 99 -c 4096 -np 1 --host 127.0.0.1 --port 8080 --chat-template-kwargs '{"enable_thinking":false}'

Confirm that --list-devices lists the RTX 5090. In another PowerShell window:

$qwenBody = @{
    messages = @(@{ role = "user"; content = "Explain gravity briefly." })
    max_tokens = 256
    temperature = 0
    stream = $false
} | ConvertTo-Json -Depth 5
$qwenReply = Invoke-RestMethod -Uri "http://127.0.0.1:8080/v1/chat/completions" -Method Post -ContentType "application/json" -Body $qwenBody -TimeoutSec 300
$qwenReply.choices[0].message.content
๐Ÿง Optional: Ubuntu under WSL2

Optional alternative, not required for native Windows. Follow Microsoft's WSL installation guide. From an administrator PowerShell, only if you want this alternative:

wsl --install -d Ubuntu

Restart if requested and complete Ubuntu's first-run account setup. Then open Ubuntu:

wsl -d Ubuntu

Inside Ubuntu, follow this device's Ubuntu/Linux instructions and create a separate Linux Python environment. Keep the Windows NVIDIA driver current; use NVIDIA's CUDA-on-WSL guidance rather than installing a Linux display driver inside WSL. Verify GPU access before loading weights.

Use either the Windows or WSL model server during the test, not both on the same port. WSL does not add VRAM. Reuse model downloads where practical, but do not reuse Windows Python environments inside Linux; Ollama stores/imports may create another copy. Native Windows remains the default path above.

Keep the service on localhost. For a remote server, use SSH port forwarding rather than opening an unauthenticated public endpoint. Start with a short context; other GPU workloads reduce available VRAM.

๐Ÿ’ป Laptop ยท 32 GB RAM / 8 GB VRAM โ€” IQ4_XS with Ollama

IQ4_XS: about 17.86 GiB, plus a 0.84-GiB vision projector and runtime memory. It does not fit fully into 8 GB VRAM; GPU + CPU/RAM offload is required. Allow about 40 GB free disk for the download and a possible second imported copy. These are planning estimates, not measured peaks. Close memory-heavy apps and use one model at a time.

The complete IQ4_XS laptop comparison is still pending. The importer was previously tested with Ollama 0.33.3; installed versions and device-specific performance are recorded separately.

๐Ÿง Ubuntu / Linux โ€” installation & launch

Install Ollama using its official Linux instructions, plus Python 3 and virtual-environment support. Start the Ollama service, or run ollama serve in a separate terminal if it is not already running. Do not start a second server on an occupied port.

mkdir -p qwen-v2-iq4
cd qwen-v2-iq4
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade huggingface_hub requests
repo="oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Ollama"
hf download "$repo" Qwen3.6-35B-v2-IQ4_XS.gguf mmproj-Qwen3.6-35B-v2-F16.gguf import_ollama.py SHA256SUMS LICENSE --local-dir .
python import_ollama.py --quant IQ4_XS
ollama run qwen3.6-35b-v2:iq4_xs --think=false

For a lighter, bounded first request instead of interactive chat:

curl http://127.0.0.1:11434/api/chat \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3.6-35b-v2:iq4_xs","messages":[{"role":"user","content":"Explain gravity briefly."}],"think":false,"stream":false,"options":{"num_ctx":2048,"num_predict":256}}'
๐ŸชŸ Windows / PowerShell โ€” native installation & launch

Install Ollama for Windows and Python 3.11. Keep the Ollama app running. The commands use a dedicated Python environment and do not change PowerShell execution policy.

New-Item -ItemType Directory -Path .\qwen-v2-iq4 -Force | Out-Null
Set-Location .\qwen-v2-iq4
py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade huggingface_hub requests
$qwenRepo = "oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Ollama"
.\.venv\Scripts\hf.exe download $qwenRepo Qwen3.6-35B-v2-IQ4_XS.gguf mmproj-Qwen3.6-35B-v2-F16.gguf import_ollama.py SHA256SUMS LICENSE --local-dir .
if ($LASTEXITCODE -ne 0) { throw "Download failed" }
.\.venv\Scripts\python.exe .\import_ollama.py --quant IQ4_XS
if ($LASTEXITCODE -ne 0) { throw "Import failed" }
ollama run qwen3.6-35b-v2:iq4_xs --think=false

For a lighter, bounded first request instead of interactive chat:

$qwenRequest = @{
    model = "qwen3.6-35b-v2:iq4_xs"
    messages = @(@{ role = "user"; content = "Explain gravity briefly." })
    think = $false
    stream = $false
    options = @{ num_ctx = 2048; num_predict = 256 }
} | ConvertTo-Json -Depth 5
$qwenReply = Invoke-RestMethod -Uri "http://127.0.0.1:11434/api/chat" -Method Post -ContentType "application/json" -Body $qwenRequest -TimeoutSec 300
$qwenReply.message.content
๐Ÿง Optional: Ubuntu under WSL2

Optional alternative, not required for native Windows. Follow Microsoft's WSL installation guide. From an administrator PowerShell, only if you want this alternative:

wsl --install -d Ubuntu

Restart if requested and complete Ubuntu's first-run account setup. Then open Ubuntu:

wsl -d Ubuntu

Inside Ubuntu, follow this device's Ubuntu/Linux instructions and create a separate Linux Python environment. Keep the Windows NVIDIA driver current; use NVIDIA's CUDA-on-WSL guidance rather than installing a Linux display driver inside WSL. Verify GPU access before loading weights.

Use either the Windows or WSL model server during the test, not both on the same port. WSL does not add VRAM. Reuse model downloads where practical, but do not reuse Windows Python environments inside Linux; Ollama stores/imports may create another copy. Native Windows remains the default path above.

The importer verifies the selected weight/projector hashes and includes both. Its interactive defaults are 4,096 context tokens and 1,024 output tokens; the API examples use 2,048 / 256. Check CPU/GPU allocation with ollama ps. No Transformers/BF16 download is needed. Ollama chat API.

๐Ÿงช Smaller experimental option โ€” IQ2_M with Ollama

IQ2_M: about 11.70 GiB, plus a 0.84-GiB vision projector and runtime memory. It does not fit fully into 8 GB VRAM; GPU + CPU/RAM offload is required. Allow about 27 GB free disk for the download and a possible second imported copy. These are planning estimates, not measured peaks. Close memory-heavy apps and use one model at a time.

Experimental low-bit option. This is IQ2_M, not Q2_K. Reasoning, code and format accuracy can suffer. Initial small Windows smoke checks have started; a full laptop quality/performance comparison is still pending.

๐Ÿง Ubuntu / Linux โ€” installation & launch

Install Ollama using its official Linux instructions, plus Python 3 and virtual-environment support. Start the Ollama service, or run ollama serve in a separate terminal if it is not already running. Do not start a second server on an occupied port.

mkdir -p qwen-v2-iq2
cd qwen-v2-iq2
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade huggingface_hub requests
repo="oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Ollama"
hf download "$repo" Qwen3.6-35B-v2-IQ2_M.gguf mmproj-Qwen3.6-35B-v2-F16.gguf import_ollama.py SHA256SUMS LICENSE --local-dir .
python import_ollama.py --quant IQ2_M
ollama run qwen3.6-35b-v2:iq2_m --think=false

For a lighter, bounded first request instead of interactive chat:

curl http://127.0.0.1:11434/api/chat \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3.6-35b-v2:iq2_m","messages":[{"role":"user","content":"Explain gravity briefly."}],"think":false,"stream":false,"options":{"num_ctx":2048,"num_predict":256}}'
๐ŸชŸ Windows / PowerShell โ€” native installation & launch

Install Ollama for Windows and Python 3.11. Keep the Ollama app running. The commands use a dedicated Python environment and do not change PowerShell execution policy.

New-Item -ItemType Directory -Path .\qwen-v2-iq2 -Force | Out-Null
Set-Location .\qwen-v2-iq2
py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade huggingface_hub requests
$qwenRepo = "oktayd/Qwen3.6-35B-v2-MoE-Ablit-Heretic-Uncensor-Hermes-MTP-Vision-Ollama"
.\.venv\Scripts\hf.exe download $qwenRepo Qwen3.6-35B-v2-IQ2_M.gguf mmproj-Qwen3.6-35B-v2-F16.gguf import_ollama.py SHA256SUMS LICENSE --local-dir .
if ($LASTEXITCODE -ne 0) { throw "Download failed" }
.\.venv\Scripts\python.exe .\import_ollama.py --quant IQ2_M
if ($LASTEXITCODE -ne 0) { throw "Import failed" }
ollama run qwen3.6-35b-v2:iq2_m --think=false

For a lighter, bounded first request instead of interactive chat:

$qwenRequest = @{
    model = "qwen3.6-35b-v2:iq2_m"
    messages = @(@{ role = "user"; content = "Explain gravity briefly." })
    think = $false
    stream = $false
    options = @{ num_ctx = 2048; num_predict = 256 }
} | ConvertTo-Json -Depth 5
$qwenReply = Invoke-RestMethod -Uri "http://127.0.0.1:11434/api/chat" -Method Post -ContentType "application/json" -Body $qwenRequest -TimeoutSec 300
$qwenReply.message.content
๐Ÿง Optional: Ubuntu under WSL2

Optional alternative, not required for native Windows. Follow Microsoft's WSL installation guide. From an administrator PowerShell, only if you want this alternative:

wsl --install -d Ubuntu

Restart if requested and complete Ubuntu's first-run account setup. Then open Ubuntu:

wsl -d Ubuntu

Inside Ubuntu, follow this device's Ubuntu/Linux instructions and create a separate Linux Python environment. Keep the Windows NVIDIA driver current; use NVIDIA's CUDA-on-WSL guidance rather than installing a Linux display driver inside WSL. Verify GPU access before loading weights.

Use either the Windows or WSL model server during the test, not both on the same port. WSL does not add VRAM. Reuse model downloads where practical, but do not reuse Windows Python environments inside Linux; Ollama stores/imports may create another copy. Native Windows remains the default path above.

The importer verifies the selected weight/projector hashes and includes both. Its interactive defaults are 4,096 context tokens and 1,024 output tokens; the API examples use 2,048 / 256. Check CPU/GPU allocation with ollama ps. No Transformers/BF16 download is needed. Ollama chat API.