File size: 3,824 Bytes
0555a7e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
---
library_name: mlx
license: apache-2.0
license_link: https://huggingface.co/DJLougen/Ornstein3.6-35B-A3B-SABER/blob/main/LICENSE
pipeline_tag: text-generation
language:
- en
tags:
- mlx
- lightning-mlx
- mtplx
- qwen3.5
- qwen3_5_moe
- mixture-of-experts
- apple-silicon
- saber
- refusal-ablation
- uncensored
base_model: DJLougen/Ornstein3.6-35B-A3B-SABER
base_model_relation: quantized
---

# Ornstein3.6-35B-A3B-SABER-4bit-MTPLX-Optimized-Speed

MLX 4-bit build of [`DJLougen/Ornstein3.6-35B-A3B-SABER`](https://huggingface.co/DJLougen/Ornstein3.6-35B-A3B-SABER) packaged for fast local serving with [`lightning-mlx`](https://github.com/samuelfaj/lightning-mlx).

The checkpoint includes an MTPLX sidecar (`mtp.safetensors`) and runtime metadata (`mtplx_runtime.json`) so `lightning-mlx` can use its Qwen3.5 MoE MTPLX serving path on Apple Silicon. Runtime metadata verified on Darwin arm64 with `mtplx_version: 0.1.0rc3`, `mtp_depth_max: 1`, `recommended_profile: sustained`.

The model is the **SABER**-ablated variant of Ornstein3.6-35B-A3B (Qwen3.5 MoE, 35B total / ~3B active per token). Refer to the [source model card](https://huggingface.co/DJLougen/Ornstein3.6-35B-A3B-SABER) for capabilities, license, and SABER details.

> **Note on MTP weights**: `mtp.safetensors` is packed from the upstream [`Qwen/Qwen3.5-35B-A3B`](https://huggingface.co/Qwen/Qwen3.5-35B-A3B) MTP module. The base model itself is the SABER fine-tune; speculative decoding acceptance rate may differ from upstream.

## Install lightning-mlx

```bash
python3 -m pip install git+https://github.com/samuelfaj/lightning-mlx.git
```

Or:

```bash
curl -fsSL https://raw.githubusercontent.com/samuelfaj/lightning-mlx/main/install.sh | bash
```

Verify:

```bash
lightning-mlx --help
```

## Serve this model

From Hugging Face:

```bash
lightning-mlx serve samuelfaj/Ornstein3.6-35B-A3B-SABER-4bit-MTPLX-Optimized-Speed
```

From a local checkout:

```bash
lightning-mlx serve /path/to/Ornstein3.6-35B-A3B-SABER-4bit-MTPLX-Optimized-Speed
```

Daemon mode:

```bash
lightning-mlx serve samuelfaj/Ornstein3.6-35B-A3B-SABER-4bit-MTPLX-Optimized-Speed --daemon
lightning-mlx status
lightning-mlx tui <PID-or-model-name>
lightning-mlx kill <PID-or-model-name>
```

## OpenAI-compatible API

```bash
curl http://localhost:8010/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "local",
    "messages": [
      {"role": "user", "content": "Write a tiny Python HTTP server."}
    ],
    "stream": true
  }'
```

## Why use lightning-mlx

`lightning-mlx` is built for local agent workloads on Apple Silicon: short streamed turns, tool calls, growing context, repeated low-latency interactions. With this checkpoint it uses the packaged MTPLX metadata and Qwen3.5 MoE serving preset instead of treating the model as a generic MLX checkpoint.

The runtime focuses on:

- OpenAI-compatible local serving
- Fast streamed chat completions
- Qwen3.5 MoE reasoning and tool-use paths
- MTPLX-style speculative decoding support
- Daemon, status, TUI, and kill controls

## Convert similar local MTPLX models

```bash
lightning-mlx convert-mtplx \
  /path/to/Model-MLX-quantized \
  --mtp-source /path/to/Model-with-mtp-tensors
```

Output is written next to the source as `<source>-MTPLX-Optimized-Speed`. Then:

```bash
lightning-mlx serve /path/to/Model-MLX-quantized-MTPLX-Optimized-Speed
```

## Use with mlx-lm

This checkpoint is also a standard MLX text-generation model:

```bash
pip install -U mlx-lm
mlx_lm.generate \
  --model samuelfaj/Ornstein3.6-35B-A3B-SABER-4bit-MTPLX-Optimized-Speed \
  --prompt "Hello" \
  --max-tokens 100
```

## Intended use

Research and red-teaming. SABER ablates refusal behaviors. Deploy behind your own policy/logging layer.

## License

Apache 2.0, inherited from the base model.