---
license: other
license_name: evrmind-free-1.0
license_link: LICENSE.md
language:
- en
library_name: llama.cpp
tags:
- deepseek
- deepseek-r1
- llama
- llama-3.1
- gguf
- 3-bit
- quantization
- evr
- evrmind
- text-generation
- reasoning
- chain-of-thought
- on-device
- bafethu
pipeline_tag: text-generation
model-index:
- name: Evrmind EVR-1 Bafethu-8b-Reasoning (DeepSeek R1 Distilled)
results:
- task:
type: text-generation
metrics:
- name: Perplexity (wikitext-2, ctx=512)
type: perplexity
value: 14.40
- name: Coherence (rep4 @ 500 tokens)
type: repetition-rate
value: 0.44
- name: Coherence (rep4 @ 1000 tokens)
type: repetition-rate
value: 1.75
---
# Evrmind EVR-1 Bafethu-8b-Reasoning (DeepSeek R1 Distilled), ~3.9 GiB
A custom 3-bit compression of DeepSeek-R1-Distill-Llama-8B that fits in **under 4 GiB** while maintaining coherent chain-of-thought reasoning at 1000+ tokens.
EVR-1 is not a standard quantization (not Q2, Q3, Q4, etc). It is a custom compression method developed by Evrmind. The compressed weights average approximately 3 bits per parameter; the total GGUF file (~3.9 GiB) includes additional metadata and structure overhead.
## Overview
- **Chain-of-thought reasoning**: The model thinks step-by-step using `...` tags before answering
- **Under 4 GiB**: Runs on laptops, desktops, and edge devices (Android via Termux also supported)
- **Coherent at 1000+ tokens**: Low repetition rate (1.75% rep4 at 1000 tokens)
## How Reasoning Works
The model uses DeepSeek R1's reasoning format. It first thinks through the problem internally, then provides a clean answer:
```
User: What is the derivative of x^3 + 2x?
To find the derivative, I apply the power rule to each term:
- d/dx(x^3) = 3x^2
- d/dx(2x) = 2
So the derivative is 3x^2 + 2.
The derivative of x^3 + 2x is **3x^2 + 2**.
```
## How to Run
Download the model file and the binary for your platform, then:
```bash
# Extract the binary
mkdir -p linux-cuda && tar xzf evrmind-linux-cuda.tar.gz -C linux-cuda
# Run (interactive chat with reasoning)
cd linux-cuda
LD_LIBRARY_PATH=. ./llama-cli -m ../evr-deepseek-r1-llama-8b-reasoning.gguf -ngl 99
# Run (single completion)
LD_LIBRARY_PATH=. ./llama-completion -m ../evr-deepseek-r1-llama-8b-reasoning.gguf -p "Your prompt here" -n 1000 -ngl 99
```
### Platform Binaries
| Platform | File | GPU Required |
|----------|------|-------------|
| Linux + NVIDIA | `evrmind-linux-cuda.tar.gz` | NVIDIA GPU (CUDA 12) |
| Linux + Any GPU | `evrmind-linux-vulkan.tar.gz` | Any Vulkan-capable GPU |
| Windows + NVIDIA | `evrmind-windows-cuda.zip` | NVIDIA GPU (CUDA 12) |
| Windows + Any GPU | `evrmind-windows-vulkan.zip` | Any Vulkan-capable GPU |
| macOS (Apple Silicon) | `evrmind-macos-metal.tar.gz` | M1/M2/M3/M4 |
| Android (Termux) | `evrmind-android-vulkan.tar.gz` | Vulkan |
> **Note:** The binaries are the same for the base, instruct, and reasoning models. You only need to download them once. Just point them at whichever GGUF you want to run.
### Flags
| Flag | Description |
|------|------------|
| `-ngl 99` | Offload all layers to GPU (recommended) |
| `-n 1000` | Generate 1000 tokens (reasoning models need more tokens for thinking) |
| `-p "..."` | Your prompt |
| `-t 8` | Number of CPU threads (for CPU layers) |
## Model Details
- **Name:** Evrmind EVR-1 Bafethu-8b-Reasoning (DeepSeek R1 Distilled)
- **Base model:** DeepSeek-R1-Distill-Llama-8B (Llama 3.1 8B architecture)
- **Size:** ~3.9 GiB (GGUF)
- **Method:** EVR-1 (Evrmind Reconstruction), a custom 3-bit compression method
- **Backends:** CUDA, Vulkan, Metal, CPU
- **Context:** Tested up to 2048 tokens; longer contexts have not been validated at 3-bit compression
- **Chat template:** DeepSeek R1 format (built-in)
## Benchmarks
### Coherence (5 continuation-style prompts, 500 and 1000 tokens each)
Average 4-gram repetition rate (lower = better):
| Model | Size | rep4 @ 500 | rep4 @ 1000 |
|-------|------|-----------|-------------|
| **EVR-1 Bafethu** | **3.93 GiB** | **0.44%** | **1.75%** |
### Perplexity
| Model | Size | Perplexity (wikitext-2, ctx=512) |
|-------|------|--------------------------------|
| DeepSeek-R1-Distill-Llama-8B Q4_K_M | 4.69 GiB | 14.39 |
| **EVR-1 Bafethu** | **3.93 GiB** | **14.40** |
## Also Available
- **[EVR-1 Maano-8b](https://huggingface.co/evrmind/evr-1-maano-8b)**, base model for text completion and creative writing
- **[EVR-1 Maano-8b-Instruct](https://huggingface.co/evrmind/evr-1-maano-8b-instruct)**, instruction-following chat model
## Intended Use
This model is intended for on-device reasoning, math, logic, and coding tasks on laptops, desktops, and edge devices where memory is constrained. An Android (Termux) build is also available. There is no iOS build.
## Limitations
- Math reasoning quality is limited by the 3-bit compression level.
- Occasional minor character-level artefacts (e.g., dropped letters) due to 3-bit compression.
- Generation quality degrades somewhat beyond 1000 tokens.
- Reasoning chains may occasionally be incomplete or circular.
- The model may identify itself as "DeepSeek-R1"; this is expected, as the underlying model was trained by DeepSeek. A system prompt can be used to override this behaviour.
- As with all heavily quantized models, generated text may contain factual inaccuracies (e.g., incorrect numbers, dates, or scientific details). Always verify factual claims independently.
## System Requirements
- **Storage:** ~4 GiB for model weights + ~50 MB for binaries
- **RAM:** 6 GiB minimum (8 GiB recommended)
- **GPU (recommended):** NVIDIA GPU with CUDA 12, Apple Silicon (M1/M2/M3/M4), or any Vulkan-capable GPU
- **CPU-only:** Supported but significantly slower
- **OS:** Linux (x86_64), macOS (Apple Silicon), Windows (x86_64), Android (Termux, ARM64)
- **Not supported:** iOS, 32-bit systems
## Safety and Responsible Use
This model inherits the capabilities and limitations of its base model (DeepSeek-R1-Distill-Llama-8B). Like all language models, it can generate incorrect, biased, or harmful content. Users should:
- Not rely on this model for factual accuracy without verification
- Not use this model to generate content that could cause harm
- Apply appropriate content filtering for any user-facing applications
- Be aware that 3-bit compression may amplify certain failure modes of the base model
- Be aware that reasoning chains may contain errors or circular logic
## Derivative Works
If you create derivative works, credit **"EVR-1 Bafethu"** in your model name and documentation. Commercial use is permitted subject to the Llama 3.1 Community License Agreement and DeepSeek MIT License.
## License
Available for personal, research, and commercial use with attribution, subject to upstream license terms. See LICENSE.md for full terms.
Built with Llama. This model is a derivative of [DeepSeek-R1-Distill-Llama-8B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-8B) (MIT License) which is based on Meta's Llama 3.1 8B. Subject to the [DeepSeek MIT License](DEEPSEEK_LICENSE.md) and the [Llama 3.1 Community License Agreement](https://www.llama.com/llama3_1/license/) in addition to the Evrmind license.
## Citation
```
@misc{evrmind2026evr1bafethu8breasoning,
title={Evrmind EVR-1 Bafethu-8b-Reasoning (DeepSeek R1 Distilled): A Custom 3-Bit Compression Method for Coherent On-Device Reasoning},
author={Evrmind},
year={2026},
url={https://huggingface.co/evrmind/evr-1-bafethu-8b-reasoning}
}
```
## Contact
- Email: hello@evrmind.io
- Issues: [GitHub](https://github.com/evrmind-uk/evr-llama/issues)