Instructions to use alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF # Run inference directly in the terminal: llama cli -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF # Run inference directly in the terminal: llama cli -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF # Run inference directly in the terminal: ./llama-cli -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
Use Docker
docker model run hf.co/alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
- LM Studio
- Jan
- Ollama
How to use alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF with Ollama:
ollama run hf.co/alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
- Unsloth Desktop
- Pi
How to use alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF with Docker Model Runner:
docker model run hf.co/alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
- Lemonade
How to use alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
Run and chat with the model
lemonade run user.Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
APEX GGUF quantizations of OS-Software/Hy-MT2-30B-A3B-uncensored-heretic, quantized by alphaZimuth.
This release is based directly on the existing uncensored / decensored model from OS-Software, rather than performing an additional uncensoring or weight-ablation process on the original Tencent model.
The source model was created from Tencent Hy-MT2-30B-A3B using Heretic v1.4.0+custom, with Arbitrary-Rank Ablation (ARA), a LoRA adapter, and row-norm preservation.
No additional uncensoring, abliteration, or behavioral weight editing is performed prior to GGUF conversion and APEX quantization.
Source Model
Source:
OS-Software/Hy-MT2-30B-A3B-uncensored-heretic
Original base model:
According to the source model authors, the decensored model was produced using:
- Heretic v1.4.0+custom
- Arbitrary-Rank Ablation (ARA)
- LoRA adapter
- Row-norm preservation
- Modified layers 18–28
The source repository should be considered the authoritative reference for the uncensoring methodology and evaluation.
Upstream Evaluation
The following results were reported by OS-Software for the source Hy-MT2-30B-A3B-uncensored-heretic model:
| Metric | Uncensored Heretic | Original Hy-MT2 |
|---|---|---|
| Keywords / Refusal Test | 0 / 100 | 100 / 100 |
| KL Divergence | 0.0276 | 0 |
The upstream evaluation used a custom mixed-language dataset.
Note
These results belong to the original full-precision source model from OS-Software.
They are included here for reference and should not be interpreted as an independent benchmark of the APEX GGUF quantizations.
Quantization can introduce small behavioral or quality differences compared with the source model.
Available Quantizations
Four APEX-I quantization tiers are currently available:
| File | Size | Description |
|---|---|---|
Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-I-Nano.gguf |
8.99 GiB | Near-limit quantization; not generally recommended |
Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-I-Mini.gguf |
10.4 GiB | A slightly higher-quality lightweight option |
Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-I-Compact.gguf |
13.0 GiB | Recommended balance between model size and quality |
Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-I-Quality.gguf |
17.7 GiB | Higher-precision variant prioritizing model quality |
Quantization Details
This repository uses APEX quantization, based on llama.cpp.
The APEX-I variants use importance-matrix-assisted quantization.
Conversion pipeline:
OS-Software/Hy-MT2-30B-A3B-uncensored-heretic
↓
convert_hf_to_gguf.py
↓
Hy-MT2-30B-A3B-uncensored-heretic-BF16.gguf
↓
Imatrix Calibration
↓
APEX Quantization
↓
APEX-I-Nano / Compact / Quality
Imatrix Calibration
The APEX quantizations in this repository were generated using an importance matrix (imatrix) calibration pass.
The calibration run reported:
Final estimate: PPL = 5.2905 +/- 0.02242
This value is the final perplexity estimate reported during the imatrix calibration process and is provided for reproducibility and reference.
Important distinction:
Uncensoring / Heretic processing → OS-Software
GGUF conversion / Imatrix calibration / APEX quantization → alphaZimuth
No additional Heretic processing or abliteration is performed in this release.
llama.cpp Compatibility
Important
Hy-MT2-30B-A3B uses the hy_v3 architecture.
Support for hy_v3 has been merged into upstream llama.cpp.
Please use a recent version of llama.cpp for inference.
Older builds may fail to load the model or may not correctly support the architecture.
Recommended Inference Parameters
The upstream Hy-MT2 documentation recommends the following parameters for the 30B-A3B model:
{
"temperature": 0.7,
"top_p": 1.0,
"top_k": -1,
"repetition_penalty": 1.0,
"max_tokens": 4096
}
About the Uncensored Version
This model has substantially reduced safety alignment compared with the original Hy-MT2 model.
As a result, it may be less likely to refuse certain requests, but it may also produce:
- inaccurate information
- harmful content
- biased content
- offensive content
- unexpected outputs
The uncensored behavior originates from the OS-Software source model and is not introduced by APEX quantization.
Users should treat model outputs as untrusted and independently verify important information.
This model is primarily intended for research, experimentation, translation research, alignment research, and local model experimentation.
Users are responsible for complying with applicable laws, regulations, licenses, and platform policies.
Previous Experimental Version
The previous release:
Hy-MT2-30B-A3B-Uncensored-v1-APEX-GGUF
used a custom directional / orthogonal ablation pipeline created specifically for that release.
Because issues were reported with the experimental version, it is now considered a legacy / experimental release.
For new deployments, use:
alphaZimuth/Hy-MT2-30B-A3B-Uncensored-Heretic-APEX-GGUF
instead.
Credits
Special thanks to:
- Tencent Hy-MT2 team for the original Hy-MT2-30B-A3B model.
- OS-Software for creating and evaluating the
Hy-MT2-30B-A3B-uncensored-hereticmodel. - p-e-w for the Heretic project.
- LocalAI for the APEX quantization framework.
- llama.cpp contributors for GGUF and
hy_v3architecture support. - The wider open-source LLM community.
APEX GGUF quantization by alphaZimuth.
License
Apache-2.0
Please also review the license and usage terms of the original Tencent Hy-MT2 model and the OS-Software derivative model.
Related Models
Original APEX release
alphaZimuth/Hy-MT2-30B-A3B-APEX-GGUF
Source uncensored model
OS-Software/Hy-MT2-30B-A3B-uncensored-heretic
Previous experimental uncensored release
- Downloads last month
- 401
We're not able to determine the quantization variants.