Instructions to use IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
Use Docker
docker model run hf.co/IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
- LM Studio
- Jan
- vLLM
How to use IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
- Ollama
How to use IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF with Ollama:
ollama run hf.co/IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF with Docker Model Runner:
docker model run hf.co/IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
- Lemonade
How to use IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
Run and chat with the model
lemonade run user.Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "IsValorum/Occamy-1.0-APEX-I-MiniPlus-V2.1-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Occamy-1.0 APEX-I-MiniPlus-V2.1 GGUF
- Optimization History & Transparency Notice
- Quick Navigation Index
- Model Files & Technical Specifications
- Surgical Tensor Quantization Map (Audited from GGUF)
- Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
- The 24GB Miracle: Full 256K Context Runs In VRAM!
- Recommended Configuration & Setup
- Optimization History & Transparency Notice
Occamy-1.0 APEX-I-MiniPlus-V2.1 GGUF
The Definitive Frontier MoE · Blistering +24 to 28+ tok/s with System RAM Offload · Full 256K Context on 24GB Workstations
THE DEFINITIVE SPECIFICATION IN THE 13–14 GB CEILING
This APEX-I-MiniPlus-V2.1 release represents the absolute technological limit of sparse Mixture-of-Experts quantization within the 13–14 GB envelope. Every single tensor of its 40 layers and 256 micro-experts has been mathematically audited to maximize reasoning precision, eliminate recurrence state drift, and prevent AVX2 CPU dequantization stalls.
DO NOT CONFUSE APEX-I-MINIPLUS WITH GENERIC COMMUNITY APEX-I-MINI!
Regardless of release version (whether V1, V2, or V2.1), NEVER confuse handcrafted APEX-I-MiniPlus builds with generic community APEX-I-Mini releases:
- Generic Community APEX-I-Mini: Uniformly compresses all core MoE experts down to aggressive 2-bit
IQ2_S(dropping below the critical quality floor), leaves the sensitive token output head unarmored at 3-bitQ3_K_M, and compresses attention projections down toQ3_K. In deep reasoning models, this triggers severe perplexity spikes, syntax errors, and broken code brackets.- Handcrafted APEX-I-MiniPlus (All Editions by IsValorum): Every single MiniPlus release—from V1 and V2 to V2.1—is a custom tensor-by-tensor architecture that preserves uncompressed
F32router gates, armors the token output head in high-precisionQ6_K, safeguards attention gates inQ8_0, and keeps core reasoning experts at or above calibrated 3-bit (IQ3_XXS/IQ3_S). Even our earlier builds vastly outperform generic community APEX recipes and flat 3-bit quants.
Optimization History & Transparency Notice
We maintain our previous releases publicly as a transparent engineering record of continuous optimization. Below is the exact evolutionary roadmap of our MiniPlus architectures:
| Specification | Core Experts (10–29) | Edge Experts (0–9, 30–39) | Shared Expert (shexp) |
Full Attention (L3, 7, 11, ...) | Attention Gates (30 Layers) | Output Head (output.weight) |
Routers (gate_inp) |
Size / Overhead | Real-World Impact |
|---|---|---|---|---|---|---|---|---|---|
| Generic APEX Mini | IQ2_S (2.50 bpw) |
Q3_K (only 5 layers) |
Q4_K / Q3_K |
Q3_K |
Compressed | Q3_K_M |
Compressed | Baseline (~12.5 GB) | Severe syntax errors, broken code indentation, high perplexity in <think>. |
| MiniPlus V2 (High Theoretical Armor) | IQ3_XXS |
IQ3_S (10 layers) |
IQ4_NL |
Q3_K |
Q8_0 |
Q6_K |
F32 |
+1.2 GB vs generic | Heavy theoretical edge envelope. In practice, virtually identical quality to V2.1 even at +160K context. Excellent for 100% VRAM offload. |
| MiniPlus V2.1 (CURRENT) | IQ3_XXS |
Q3_K (10 layers) |
Q5_K (All 40 layers) |
Q4_K (q/k/v) + Q6_K (output) |
Q8_0 |
Q6_K |
F32 |
< 100 MB extra over V2 (~13.74 GiB total) | Zero AVX2 CPU stalls; +24 to 28+ tok/s streaming when offloading bulk of the model to system RAM (DDR4/DDR5). Only ~100 MB overhead over V2 (completely negligible in RAM). If 100% in VRAM, both perform identically. |
Architecture & Edition Guide — Choosing Between Editions
- Full GPU VRAM Offload (24GB+ VRAM,
-ngl 99): Both V2 and V2.1 run blistering fast on GPU tensor cores with virtually identical top-tier quality.- In Practical Long-Context (+160K tokens): Although V2 provides higher theoretical protection on paper, real-world benchmarks show virtually zero perceptible quality difference compared to V2.1 even across deep +160K contexts.
- System RAM Streaming Specialist (DDR4/DDR5 & Massive Context): V2.1 is specially engineered to run either partially or entirely out of system RAM across large or full (+160k to 256k) context windows. By replacing non-linear codebooks with linear SIMD-optimized
Q3_Kedge experts and upgrading shared foundation experts toQ5_Kacross all 40 layers, AVX2 CPU dequantization stalls are completely eliminated (+24 to 28+ tok/s). Depending on your processor architecture and memory bandwidth (dual-channel DDR4 or high-speed DDR5 6000+ MT/s), streaming generation speeds in system RAM can approach speeds remarkably close to full VRAM execution, allowing the dedicatedQ8_0multimodal vision projector (mmproj) to be loaded explicitly in GPU VRAM for instant, zero-latency visual document parsing and OCR while the vast language weights stream economically from system RAM. The ~100 MB difference over V2 is completely negligible when running in system RAM.Both editions are handcrafted and vastly outperform flat 3-bit quants and generic community APEX-I-Mini releases.
Prefer high theoretical edge layer protection on paper? Explore the Occamy-1.0 MiniPlus V2 Edition.
Quick Navigation Index
- Model Files & Technical Specifications
- Surgical Tensor Quantization Map (Audited from GGUF)
- Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
- The 24GB Miracle: Full 256K Context Runs In VRAM!
- Recommended Configuration & Setup
Model Files & Technical Specifications
| File Name | File Size | Memory Footprint | BPW | Description |
|---|---|---|---|---|
Occamy-1.0.APEX-I-MiniPlus-V2.1.gguf |
14.75 GB (13.74 GiB) |
13.74 GiB |
3.40 BPW | Core hybrid linear attention, math, logic & multimodal vision MoE |
mmproj-Accio-Lab_occamy-1.0-Q8_0.gguf |
614 MB (585 MiB) |
585 MiB |
8.50 BPW | Dedicated Q8_0 vision projector for optical document parsing |
- Base Model: Accio-Lab/occamy-1.0
- Parameters: 35.2B total (approx. 2.6B to 3.2B active per token)
- Architecture: 40 layers, 256 micro-experts (8 active per token) + hybrid linear attention / DeltaNet recurrent layers
- Context Length: 262,144 tokens (native 256K)
Surgical Tensor Quantization Map (Audited from GGUF)
The exact tensor breakdown below has been verified directly from the compiled binary weights:
| Layer Group | Sub-Component / Tensor | Qty | Precision | Engineering Rationale |
|---|---|---|---|---|
| Global Output Head | output.weight |
1 | Q6_K |
Preserves near-FP16 token classification; eliminates syntax errors, bracket drops, and hallucinations. |
| Global Embeddings | token_embd.weight |
1 | Q4_K |
High-fidelity vocabulary embedding representation. |
| All Normalizations | output_norm, attn_*_norm, ssm_norm |
171 | F32 |
100% uncompressed numerical stability across all 40 layers. |
| Expert Routers | blk.*.ffn_gate_inp, ffn_gate_inp_shexp |
80 | F32 |
100% uncompressed routing fidelity across 256 micro-experts; zero router drift. |
| Attention Gates | blk.*.attn_gate.weight (30 Hybrid Layers) |
30 | Q8_0 |
High-precision attention gating across hybrid DeltaNet recurrence layers; eliminates crosstalk. |
| Shared Foundation Experts | blk.*.ffn_{gate,down,up}_shexp (All 40 Layers) |
120 | Q5_K |
Foundation knowledge backbone active on 100% of tokens; protected in high-precision linear Q5_K. |
| Periodic Full Attention | blk.{3,7,11,...}.attn_q/k/v (10 Anchor Layers) |
30 | Q4_K |
Full quadratic attention anchor checkpoints for deep needle-in-a-haystack retrieval. |
| Periodic Full Attention | blk.{3,7,11,...}.attn_output (10 Anchor Layers) |
10 | Q6_K |
Armored attention output projection over deep context. |
| Recurrent SSM Scales | blk.*.ssm_alpha, ssm_a, ssm_conv1d, ssm_dt |
120 | F32 |
Guarded in uncompressed FP32 to prevent DeltaNet recurrent state drift. |
| Linear Attention & SSM | blk.*.attn_qkv, ssm_beta, ssm_out |
90 | Q3_K |
Linear AVX2 execution; zero SIMD CPU stalls during system RAM streaming. |
| Edge MoE Experts | Layers 0–9 & 30–39 (ffn_*_exps) |
60 | Q3_K |
Linear SIMD execution; enables +24 to 28+ tok/s streaming under system RAM offload. |
| Core MoE Experts | Layers 10–29 (ffn_*_exps) |
60 | IQ3_XXS |
Calibrated with importance matrix (imatrix) for maximum compactness in deep layers. |
Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
Empirically verified in Unsloth Studio & llama.cpp:
| Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |
|---|---|---|---|---|
| NVIDIA RTX 5080 / 5090 (Blackwell) | Full GPU (-ngl 99) |
120 – 145+ tok/s | 2,800 – 3,900+ tok/s | Blistering throughput on GDDR7 bandwidth |
| NVIDIA RTX 4090 (24GB GDDR6X) | Full GPU (-ngl 99) |
90 – 115+ tok/s | 2,000 – 2,800+ tok/s | Linear attention layers slash prefill latency |
| NVIDIA RTX 3090 (24GB GDDR6) | Full GPU (-ngl 99) |
72 – 88+ tok/s | 1,500 – 2,200+ tok/s | Full 256k native window in VRAM |
| Workstation / Laptop (DDR4 / DDR5 RAM) | Hybrid Offload (Few layers in VRAM) | 24.25 – 28.37 tok/s | 385 – 410+ tok/s | Zero AVX2 CPU stalls; fast streaming from system RAM |
- Aggressive Hybrid Offload Profile: Sustained 24.25 to 28.37 tok/s generation with reasoning enabled, even when only ~4.2 GB VRAM is available and the rest of the 13.74 GiB model streams from system RAM.
Empirical Testbed Architecture & Desktop/Server Scaling
- Empirical Benchmark Hardware: The hybrid offload and system RAM streaming figures documented above (sustaining 24.25 to 28.37 tok/s) were measured on a consumer laptop powered by an Intel 12th Gen Alder Lake architecture featuring a hybrid design of Performance Cores (P-Cores) and Efficient Cores (E-Cores) paired with dual-channel system RAM and constrained laptop power/thermal envelopes.
- Thread Scheduling & E-Core Contention: In hybrid architectures like Alder Lake, OS thread scheduling across background E-Cores and lower single-core mobile power limits introduce memory bandwidth and thread synchronization overhead during CPU dequantization.
- Dramatic Scaling on Higher-End Processors: When running on desktop or server processors (such as modern AMD Ryzen 7000 / 9000 Zen 4/5 series or high-TDP Intel desktop platforms with dedicated performance cores, large L3 caches, and high-bandwidth dual- or quad-channel DDR5 running at 6000+ MT/s), streaming generation speeds and prefill throughput will scale dramatically higher, substantially exceeding these measured mobile numbers.
The 24GB Miracle: Full 256K Context Runs In VRAM!
Occamy-1.0 APEX-I-MiniPlus-V2.1 fits the entire 256K context window within 24GB VRAM:
| Context Length | Model Weights (Est.) | KV Cache (q8_0, 4 slots) | Compute Buffers | Total GPU VRAM (Est.) | Feasibility |
|---|---|---|---|---|---|
| 32,768 (32k) | 13.74 GiB |
0.58 GiB |
1.80 GiB |
16.12 GiB |
Full offload on 24GB; partial on 16GB |
| 65,536 (64k) | 13.74 GiB |
0.92 GiB |
1.95 GiB |
16.61 GiB |
Effortless fit on 24GB GPUs |
| 131,072 (128k) | 13.74 GiB |
1.58 GiB |
2.22 GiB |
17.54 GiB |
Effortless fit on 24GB GPUs |
| 262,144 (256k) | 13.74 GiB |
2.92 GiB |
2.80 GiB |
19.46 GiB |
FULL 256K NATIVE IN VRAM! |
Note: Leaves comfortable headroom for display drivers and compute buffers on standard 24GB GPUs (RTX 3090, RTX 4090, RTX 5090).
Recommended Configuration & Setup
llama-server.exe \
-m Occamy-1.0.APEX-I-MiniPlus-V2.1.gguf \
--port 8080 \
--parallel 4 \
--flash-attn on \
--fit on \
-c 104960 \
--cache-type-k q8_0 \
--cache-type-v q8_0
- Downloads last month
- 660
We're not able to determine the quantization variants.